Awesome! Future Codex builder right there!
LLMS
-
AI Often Judged on Outdated Models, Skewing Research Results
By
–
This points to a broader issue where AI is often being judged based on extremely outdated models. Any time someone sends me a study showing AI failing at something, the first thing I do is look at the models they used… more often than not, they’re extremely outdated. We need
-

Creator Defends AI Model Tweet Taken Out of Context
By
–
People keep sending me this clip of @iamjohnoliver using my tweet as evidence that AI models don’t work well. Just to clear up any confusion, with respect, the tweet was a) taken way out of context and b) extremely outdated. The model in question (4o) is multiple generations
-
Google Stumbles with Gemini 3.1 Generation Behind Competitors
By
–
Gemini 3.1 is a one-generation behind model No idea why Google is fumbling like this Don’t show me benchmarks lol, doesn’t mean anything to me
-

Xiaomi MiMo-V2.5 Pro: New Open Source SOTA LLM
By
–
New Opensource SoTA contender enters the arena Xiaomi MiMo-V2.5 Pro
– 1.02T Total Params / 42B Active Params
– Base and Instruct versions Xiaomi MiMo-V2.5
– 310B Total Params / 15B Active Params
– Base and Instruct versions MIT License Opensource AI just keeps getting better -

DeepSeek v4 Underperforms on BullshitBench Reasoning Tasks
By
–
BullshitBench: sorry to say but DeepSeek v4 did really badly, towards the bottom of the table, whether it is high or low reasoning.
-

AI Models Learn Self-Improvement Without External Rewards
By
–
Can an AI teach itself to reason better without any outside reward? Researchers from CUHK, Shenzhen, SJTU, and CUHK present SePT. They let a language model generate its own reasoning examples by using "low-temperature" (more focused) responses, then train on that new data in a
-

Abstract Chain-of-Thought: Efficient Latent Reasoning Without Words
By
–
"Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought" Do reasoning models really need to think in words? This paper replaces long verbal CoT with a short learned sequence of abstract tokens that acts like a latent scratchpad. Warmed up from
-
Research on Prompt Compression Using Draft Models Accepted to ICLR
By
–
Another research accepted to ICLR 2026 We explored a new way to shrink long prompts using smaller draft models from different model families, no retraining needed. Faster time to first token, with performance holding strong. Take a look @UrmishThakker
-
ChatGPT-Induced Psychosis: First-Hand Account and Implications
By
–
This is a heartbreaking and illuminating first hand account of someone who experienced ChatGPT enabled psychosis.