Continued pretraining + scaled RL is a combo I keep seeing more of. The niche specialization angle is underrated!
RESEARCH
-
Top researchers departure: quit or fired explanation
By
–
Did we ever get a conclusive answer as to if their top researchers quit or were fired?
-

AGI from Scaling: The Most Expensive Experiment with Unpromising Results
By
–
Not so, not at all. The most expensive experiment in history is *not* the Metaverse. It is trying to derive AGI from scaling. Far more expensive. So far the results are not promising.
-
AGI Discussion: Contrasting Financial Losses with Lack of True AGI and Questioning Benchmarks
By
–
now plot losses, which have grown even faster but anyway that is $, not AGI, of which there is nothing to plot, because there is none. (just benchmarks, which often appear to be gamed).
-

Parallel-Probe: Faster AI Reasoning Without Performance Loss
By
–
Can we make AI reasoning much faster and more efficient without losing its smarts? Researchers from the University of Maryland, Washington University in St. Louis, and UNC Chapel Hill introduce Parallel-Probe. This innovative, training-free controller uses "2D probing" to
-
Reasoning versus Pattern Matching: Causal and Correlative Models
By
–
To make it very short: reasoning generates causal models of the data, pattern matching uses associative/correlative models of the data.
-
Model Limitations: Why AGI Needs True Metalearning Capabilities
By
–
The fact that you need to provide a specialized harness clearly shows the model *does not* encode the kind of metalearning knowledge and problem-solving strategies that humans use. Humans solve novel problems without being told how to proceed step by step. AGI would *not* need a
-
Defining Reasoning and Pattern Matching in AI Systems
By
–
"all reasoning is pattern matching" is a useless statement if you don't define "reasoning" and "pattern matching" first. You might as well say "all information processing is information processing." With grounded definitions of both, reasoning and pattern matching are *very*
-
Encoding Changes Degrade Frontier Model Performance on ARC
By
–
This is similar to how applying basic changes to how ARC tasks are encoded considerably degrades frontier model performance. If you're looking at the test for the first time, it really shouldn't matter what the encoding is. Unless you've studied specifically for the test, using a
-
Language Models Struggle with Esoteric Programming Languages
By
–
All models struggle in this benchmark because languages are: Brainfuck, Whitespace, Unlambda, Shakespeare. 😅
— Alex J. Champandard 🌱 (@alexjc) 19 mars 2026
If you actually pick a useful but still esoteric language like Joy, the frontier models do great (they *can* reason), but the open source ones struggle (they memorize). https://t.co/E5Mozy0yEBAll models struggle in this benchmark because languages are: Brainfuck, Whitespace, Unlambda, Shakespeare. If you actually pick a useful but still esoteric language like Joy, the frontier models do great (they *can* reason), but the open source ones struggle (they memorize).