6/ The Illusion of State in State-Space Models – investigates the expressive power of state space models (SSMs) and reveals that it is limited similar to transformers in that SSMs cannot express computation outside the complexity class 𝖳𝖢^0…
@dair_ai
-
Chinchilla Scaling: Replication Attempt Challenges Hoffmann et al.
By
–
3/ Chinchilla Scaling: A replication attempt – attempts to replicate the third estimation procedure of the compute-optimal scaling law proposed in Hoffmann et al. (2022) (i.e., Chinchilla scaling); finds that “the reported estimates are inconsistent with their first two
-
Llama 3 Family: 8B and 70B Models Released with Strong Performance
By
–
1/ Llama 3 – a family of LLMs that include 8B and 70B pretrained and instruction-tuned models; Llama 3 8B outperforms Gemma 7B and Mistral 7B Instruct; Llama 3 70 broadly outperforms Gemini Pro 1.5 and Claude 3 Sonnet.https://t.co/xqDJy0f0Vs
— DAIR.AI (@dair_ai) 21 avril 20241/ Llama 3 – a family of LLMs that include 8B and 70B pretrained and instruction-tuned models; Llama 3 8B outperforms Gemma 7B and Mistral 7B Instruct; Llama 3 70 broadly outperforms Gemini Pro 1.5 and Claude 3 Sonnet.
-
Top ML Papers Week: Llama 3, Mixtral, RAG Advances
By
–
The Top ML Papers of the Week (April 15 – April 21): – Llama 3
– Mixtral 8x22B
– A Survey on RAG
– How Faithful are RAG Models?
– Emerging AI Agent Architectures
– Chinchilla Scaling: A replication attempt
… -

NLP Cross-Field Engagement Declined Significantly Since 1980
By
–
10/ The Influence Between NLP and Other Fields – aims to quantify the degree of influence between 23 fields of study and NLP; the cross-field engagement of NLP has declined from 0.58 in 1980 to 0.31 in 2022…
-

Aligning LLMs to Quote from Pre-Training Data
By
–
9/ Aligning LLMs to Quote from Pre-Training Data – proposes techniques to align LLMs to leverage memorized information quotes directly from pre-training data.
-

Knowledge Capacity Scaling Laws in Language Models
By
–
8/ The Physics of Language Models – investigates knowledge capacity scaling laws where it evaluates a model’s capability via loss or benchmarks, to estimate the number of knowledge bits a model stores.
-

Multilingual LLMs Survey: Methods, Taxonomy and Research Frontiers
By
–
7/ Overview of Multilingual LLMs – a survey on multilingual LLMs including a thorough review of methods, a taxonomy, emerging frontiers, challenges, and resources to advance research.
-
Iterative Self-Revision and Search for LLM Reasoning Tasks
By
–
6/ Reasoning with Intermediate Revision and Search – presents an approach for general reasoning and search on tasks that can be decomposed into components; incorporates iterative self-revision capabilities and allows an LLM to build an interwoven network of thoughts.
-
Google DeepMind Synthetic Data Research Best Practices Overview
By
–
5/ Best Practices and Lessons on Synthetic Data – an overview by Google DeepMind on synthetic data research, covering applications, challenges, and future directions
