Trending on alphaXiv (2/6): The most comprehensive (and refreshingly clear) study on how LLMs learn to reason. A systematic investigation reveals key ingredients: SFT initialization helps, reward shaping stabilizes training, filtered verifiable rewards improve generalization,
LLMS
-
Cerebras Powers Mistral Le Chat at 1,100 Tokens Per Second
By
–
Cerebras is proud to be powering the new Le Chat! We enable Flash Answers to run at over 1,100 tokens/s – 10x faster than ChatGPT 4o, Sonnet 3.5, and DeepSeek R1. Try Le Chat: https://
chat.mistral.ai/chat
Learn more: -

OpenAI Updates o3-mini Chain of Thought Reasoning Capabilities
By
–
Updated chain of thought in OpenAI o3-mini for free and paid users, and in o3-mini-high for paid users.
-
DeepSeek-R1 RL Training Simplicity Versus Process Reward Models
By
–
From an open-research point of view, probably the greatest thing about DeepSeek–R1 is how its RL training technique appears so straightforward and simple in comparison to the cumbersome approaches people were starting to think necessary for learning reasoning like Process Reward
-

Evolution of RAG Architectures: From Naive to Multi-Agent Systems
By
–
Retrieval-Augmented Generation (RAG) is evolving fast! This infographic breaks down different RAG architectures – from Naïve RAG to Multi-Agent RAG – enhancing retrieval, ranking, and generation for AI applications. What’s the future of RAG? Multi-modal, agentic, and
-

Knowledge Distillation Gains Spotlight Through DeepSeek Research
By
–
Distillation has been on the news (!) due to @deepseek_ai
. The paper https://
arxiv.org/abs/1503.02531 was actually rejected from NeurIPS 2014 due to lack of novelty (true-ish), and lack of impact . Thanks reviewer#2 (literally), and thanks for @arxiv
! @geoffreyhinton @JeffDean -
Schmidhuber’s 1991 RNN History Compression Method for Long Context
By
–
PS @SchmidhuberAI in 1991 published something related to compressing history length in RNNs, which may be very relevant for long context, an important topic of research today.
-
Cursor Uses Gemini 2 Flash for arXiv Paper Analysis
By
–
We used Gemini 2 Flash to build Cursor for arXiv papers
— alphaXiv (@askalphaxiv) 6 février 2025
Highlight any section of a paper to ask questions and “@” other papers to quickly add to context and compare results, benchmarks, etc. pic.twitter.com/2KKTuDuzRhWe used Gemini 2 Flash to build Cursor for arXiv papers Highlight any section of a paper to ask questions and “@” other papers to quickly add to context and compare results, benchmarks, etc.
-
Mistral Large Chat Now Achieves 1000+ Tokens Per Second Speed
By
–
Le chat now runs Mistral Large at 1000+ tokens/s !https://t.co/joPeKtQg7C https://t.co/9mCbx3sMoM
— Guillaume Lample @ NeurIPS 2024 (@GuillaumeLample) 6 février 2025Le chat now runs Mistral Large at 1000+ tokens/s ! https://
chat.mistral.ai
