Why Can't Transformers Learn Multiplication? This paper found that plain training never builds long-range links of multiplications. So by adding a new auxiliary loss that predicts the “running sum”, it enables the model to successfully learn multi-digit multiplication!
@askalphaxiv
-
NotebookLM Transforms ArXiv Papers Into Engaging AI Conversations
By
–
Introducing NotebookLM for arXiv papers 🚀
— alphaXiv (@askalphaxiv) 15 octobre 2025
Transform dense AI research into an engaging conversation
With context across thousands of related papers, it captures motivations, draws connections to SOTA, and explains key insights like a professor who's read the entire field pic.twitter.com/A37xmW1G4GIntroducing NotebookLM for arXiv papers Transform dense AI research into an engaging conversation With context across thousands of related papers, it captures motivations, draws connections to SOTA, and explains key insights like a professor who's read the entire field
-

Stanford Research: Competitive Pressure Breaks LLM Alignment
By
–
Fascinating alignment research from Stanford When optimized for sales, elections or social media, LLMs tend to push towards deception and divisive rhetoric This shows that competitive pressure alone can break alignment, creating what the researchers call Moloch’s Bargain
-

Automatic Language Detection and Title Generation for arXiv Articles
By
–
https://alphaxiv.org/pdf/2509.22358
-

Meta’s Stochastic Activation Achieves 90% Sparsity LLM Speedup
By
–
First ever activation function swapping strategy just dropped for LLMs! Meta’s new Stochastic Activation proposes a new method which randomly selects between non-linear functions (ReLU & SiLU) in the FFN of LLM, achieving an activation sparsity of 90% with 1.65x CPU speedups!
-

Multiple RL Trajectories Improve Agent Strategy Diversity
By
–
We have always been using 1 RL trajectory for training. Why not use more? In this research: Polychromic Objectives for RL, they train their model on sets of trajectories and reward both success & diversity, so the agent keeps MULTIPLE strategies alive all at once.
-
Google DeepMind Dreamer 4 Mines Minecraft Using World Models
By
–
Crazy paper from Google DeepMind
— alphaXiv (@askalphaxiv) 7 octobre 2025
Dreamer 4 can mine Minecraft diamonds without ever touching the real game
It trained entirely inside its own world model, kinda like an “imaginary world”, showing agents can tackle long horizon tasks safely by just learning from videos! pic.twitter.com/5cG9oYyt51Crazy paper from Google DeepMind Dreamer 4 can mine Minecraft diamonds without ever touching the real game It trained entirely inside its own world model, kinda like an “imaginary world”, showing agents can tackle long horizon tasks safely by just learning from videos!
-

Bytedance GRPO Paper Challenges Optimization Approach Efficiency
By
–
Following an intense weekend of GRPO discussion, Bytedance put out a paper on why GRPO is NOT optimal! Similar to the classic knapsack problem, they suggest that each exploration has a “value” and “cost”, which needs to be adaptively distributed. This yields 20-40% more signal
-

Scaling Parallel Agents for Computer Use Tasks
By
–
what if we stopped betting everything on one agent rollout? "The Unreasonable Effectiveness of Scaling Agents for Computer Use" Generates multiple trajectories in parallel & selects the best using "behavior narratives" 69.9% on OSWorld, nearly matching human-level 72%
-
New AI Evals Community Launches Research Discussion Platform
By
–
If you're interested in research and developments on AI evals, join our new AI Evals community! We'll be working with @valsai to host discussion and talks from the researchers behind influential papers and benchmarks. For the the first talk, the community is hosting
