6/ The Illusion of State in State-Space Models – investigates the expressive power of state space models (SSMs) and reveals that it is limited similar to transformers in that SSMs cannot express computation outside the complexity class π³π’^0…
LLMS
-
Chinchilla Scaling: Replication Attempt Challenges Hoffmann et al.
By
–
3/ Chinchilla Scaling: A replication attempt – attempts to replicate the third estimation procedure of the compute-optimal scaling law proposed in Hoffmann et al. (2022) (i.e., Chinchilla scaling); finds that βthe reported estimates are inconsistent with their first two
-
Top ML Papers Week: Llama 3, Mixtral, RAG Advances
By
–
The Top ML Papers of the Week (April 15 – April 21): – Llama 3
– Mixtral 8x22B
– A Survey on RAG
– How Faithful are RAG Models?
– Emerging AI Agent Architectures
– Chinchilla Scaling: A replication attempt
… -
Llama 3 Family: 8B and 70B Models Released with Strong Performance
By
–
1/ Llama 3 – a family of LLMs that include 8B and 70B pretrained and instruction-tuned models; Llama 3 8B outperforms Gemma 7B and Mistral 7B Instruct; Llama 3 70 broadly outperforms Gemini Pro 1.5 and Claude 3 Sonnet.https://t.co/xqDJy0f0Vs
— DAIR.AI (@dair_ai) 21 avril 20241/ Llama 3 – a family of LLMs that include 8B and 70B pretrained and instruction-tuned models; Llama 3 8B outperforms Gemma 7B and Mistral 7B Instruct; Llama 3 70 broadly outperforms Gemini Pro 1.5 and Claude 3 Sonnet.
-
Anticipating Code Interpreter Feature on Claude 3
By
–
Really looking forward to Code Interpreter on Claude 3
-

Building AI-powered blogging platform with Next.js and Langchain
By
–
Build an AI-powered blogging platform (Next.js, Langchain & CopilotKit) In this article, you will learn how to build an AI-powered blogging platform that can search the web and research any topic for a blog article https://
dev.to/copilotkit/how
-to-build-an-ai-powered-blogging-platform-nextjs-langchain-supabase-1hdp
β¦ -
Llama 3 token vocabulary expansion optimization benefits
By
–
I don't think I've talked about going beyond tokenization? Still seems like a sensible approach to me – Llama 3 upped their token vocabulary from 32,000 to 128,000 which gave them some optimization benefits
-
Base Models Predict Individuals, Not Average Humans
By
–
Also, even the current base models are in-principle being trained to predict every individual human on the Internet, not to imitate an average human on the Internet, and performance at that task would not saturate at human-level intelligence. https://
lesswrong.com/posts/nH4c3Q9t
9F3nJ7y8W/gpts-are-predictors-not-imitators
β¦ -

Dataset breakthrough potentially as impactful as GPT-4 model
By
–
Datasets might be more impactful than models at this point and this may be the GPT4 of datasets. Courtesy of the amazing Guilherme who trained Falcon & the @huggingface team!
-
Llama3 Models Analysis on Hugging Face Repository
By
–
yup it looks like some of them are not llama3. probably more ~500 are llama3 based (everything older than that can't be llama3: https://
huggingface.co/models?p=17&so
rt=modified&search=llama3
β¦)
