Very cool to see that in the past year, @huggingface has become the platform for leaderboards, benchmark and evaluations, which are very critical pieces of the AI building process. You can find a lot of these like the chatbot Arena, the openLLM leaderboard, the bigcode
LLMS
-

Google AI Secrets Stolen: Top AI Stories Today
By
–
Top stories in AI today: -Google engineer steals AI secrets
-Inflection upgrade nears GPT-4
-Generate an AI song with just a prompt
-Researchers create self-spreading AI malware
-6 new AI tools & 4 new AI jobs Read more: http://
therundown.ai/p/googles-ai-s
ecrets-stolen
… -

Claude-3 Opus beats GPT-4, GPT-4 Turbo stays top
By
–
The Chatbot Arena was just updated with the ELO rankings for the new Claude-3 models! – 𝗖𝗹𝗮𝘂𝗱𝗲-𝟯 𝙊𝙥𝙪𝙨 𝗯𝗲𝗮𝘁𝘀 𝗚𝗣𝗧-𝟰!
It's the first LLM to beat GPT-4 since its release 1 year ago. – GPT-4 Turbo stays on top with a comfortable lead of ~20 points. -
01AI Yi Successfully Upscales Yi-6B to Yi-9B Model
By
–
It's super exciting to see more success using DUS. @01AI_Yi Great job! Well done! "Method Following the methodology outlined by Kim et al. [38], our goal is to upscale our Yi-6B base model, which has 32 layers, to a 9B model named the Yi-9B base model, featuring 48
-
Find LLM leaderboards: Aymeric Roucher shares ultimate collection by C. Le Fourrier
By
–
Are you trying to find good leaderboards to compare LLMs? @clefourrier is building the ultimate collection here:
-

Kirk Borne Receives Updated 3rd Edition on Transformers and Generative AI
By
–
WOW! I just received (minutes ago) my copy of the updated 3rd Edition of this book: http://
amzn.to/3TkBknI
"Transformers for Natural Language Processing and Computer Vision: Explore #GenerativeAI and Large Language Models #LLMs with Hugging Face, ChatGPT, GPT-4V, and DALL-E 3" -
Testing AI Model Performance with Real Out-of-Distribution Data
By
–
Can you run a needle in a haystack test with this! Interested in how those do with real data that's definitely not in the training set.
-
Claude 3 Opus Outperforms GPT-4 Turbo in Advanced Reasoning
By
–
Excited to see @AnthropicAI Claude 3 Opus out-perform @OpenAI GPT4Turbo in our evals and reasoning/code gen tasks. Having multiple options for models with advanced reasoning abilities is going to be a significant boon for AI applications!
-

Claude 3 Opus Outperforms GPT-4 in Natural Writing
By
–
Anthropic unveiled Claude 3, with the top-tier ‘Opus’ version outperforming GPT-4. I've personally been testing Opus, and it feels very powerful, especially in use cases such as natural writing and editing writing. Learn more: https://
therundown.ai/p/claude-dethr
ones-gpt4
… -
AI Giants Clash: GPT Competitors, Legal Battles, and Robotics
By
–
It's been an insane week for AI. Two GPT-4 level competitors, Google's AI espionage, Elon suing OpenAI, Midjourney drama, open-source robotics, and more. Here's everything that you need to know:
