I feel like OpenAI is the only lab that really nailed reasoning. DeepSeek was probably the closest, but need to see a frontier model from them to be sure. Gemini's reasoning was quite weird and all over the place. Claude's reasoning never used to matter (non thinking were
RESEARCH
-
Iterative Testing Reveals Design Flaws Better Than Pure Thinking
By
–
You cannot think your way to a perfect design. Only building and testing, over many iterations, can reveal the flaws in your mental model and provide the feedback you need to create the best design possible.
-

Technological Revolution Labor Market Effects Economists Analysis
By
–
Dario is wrong.
— Yann LeCun (@ylecun) 18 avril 2026
He knows absolutely nothing about the effects of technological revolutions on the labor market.
Don't listen to him, Sam, Yoshua, Geoff, or me on this topic.
Listen to economists who have spent their career studying this, like @Ph_Aghion , @erikbryn ,… https://t.co/PI3q8ZsobSDario is wrong.
He knows absolutely nothing about the effects of technological revolutions on the labor market. Don't listen to him, Sam, Yoshua, Geoff, or me on this topic.
Listen to economists who have spent their career studying this, like @Ph_Aghion , @erikbryn , -

Master Any LLM: Comprehensive Guide to Large Language Models
By
–
Master Any #LLM
by @ingliguori #GenerativeAI #ArtificialIntelligence #MachineLearning #MI -
Better terminology for multidimensional arrays in machine learning
By
–
Do you have a better name than "multidimensional array" though? The name tensor is convenient, if mathematically inaccurate.
-
Tensor Engine History: From Bell Labs to PyTorch
By
–
The tensor engine was first implemented inside SN3 (before it was called Lush) in 1992 at Bell Labs by Léon Bottom and me.
The naming convention has survived to this day in PyTorch and other libraries. The naming of the tensor operations was reused in EBlearn (C++ deep learning -

Apple Attention to Mamba Cross-Architecture Distillation Technique
By
–
NEW paper from Apple. Interesting idea: "Attention to Mamba". The paper introduces a two-stage recipe for cross-architecture distillation from Transformers into Mamba. Naive distillation collapses teacher performance. Their trick: first distill the transformer into a
-
Grok 4.4 and 4.5 Release Timeline Announced
By
–
Supplemental training has been added to 4.3. Grok 4.4 will be twice the size (1T) with training data through early April. Probably ready for release in early May. Grok 4.5 will be 1.5T and hopefully out by late May.
-

ABC of AI Use Cases: Comprehensive Overview
By
–
ABC of #AI Use Cases
by @ingliguori #ArtificialIntelligence #MachineLearning #ML #DL -
Sophisticated Views on AI Models and Their Relative Strengths
By
–
i really appreciate when people develop sophisticated views about our models and their relative strengths! it is certainly a labor of love