The opacity of neural networks has infected the papers about them.
RESEARCH
-

Stochastic Tiny Recursive Models améliorent PPBench
By
–
“Probabilistic Tiny Recursive Model” This paper makes Tiny Recursive Models stochastic at test time by adding Gaussian noise, running parallel rollouts, and using the existing Q head to pick the best answer. With no retraining and no task-specific tricks, its PPBench jumps from
-
LLMs: symbolic AI triumph with more noise
By
–
Looked at this way, LLMs are the triumph of symbolic AI, just with more noise.
-
AI model quality and price evolution over time
By
–
This is elite data – how the pareto frontier moved over time. Took a lot of effort to get right. Huge shift in the model quality & price in the last 3 years. https://t.co/CtafQ4Gd66
— Peter Gostev (@petergostev) 21 mai 2026This is elite data – how the pareto frontier moved over time. Took a lot of effort to get right. Huge shift in the model quality & price in the last 3 years.
-
Testing Gemma 4 and Qwen with SQL generation quality comparison
By
–
I've tried it with Gemma 4 and Qwen 3.5/3.6 so far, works well with the >4B models, sometimes works with 4B but they're more likely to mess up the SQL
-
Questioning the Generalization Capabilities of AI in Mathematics
By
–
can’t believe people assume that success on highly verifiable problems in math (where we don’t even know how many tests were performed and how many might have failed) iautomatically generalize to everything else when there is not a shred of evidence that they do.
-

Karpathy launches team for Claude-powered recursive self-improvement
By
–
Karpathy will help launch a new team focused on using Claude itself to accelerate pretraining research. Its team is focused recursive self improvement.
-

alphaXiv s’associe à l’ACM CAIS
By
–
alphaXiv ACM CAIS We’re excited to announce our partnership with the ACM Conference on AI and Agentic Systems! alphaXiv will serve as the complimentary research hub for accepted papers and system demos, helping the community discover, understand, and build on top of the
-
Active Graph video: flipped agent architecture, blackboard system, self-improvement
By
–
longer form video on Active Graph [7 min 22 sec]
— Yohei (@yoheinakajima) 21 mai 2026
– flip the agent architecture
– 1970s blackboard system
– rollback, fork, diff agent runs
– experience + behaviors + beliefs = you
– behaviors can write behaviors (self-improvement)
github: https://t.co/8ENyNt4gYD pic.twitter.com/WeN2D60bdjlonger form video on Active Graph [7 min 22 sec] – flip the agent architecture
– 1970s blackboard system
– rollback, fork, diff agent runs
– experience + behaviors + beliefs = you
– behaviors can write behaviors (self-improvement) github: https://
github.com/yoheinakajima/
activegraph
… -
Using AI for Peer Review Alongside Humans
By
–
The implication is that you should be using AI for peer review, but combine it with humans, though AI reviewers keep getting better and humans don't. Paper: