What if LLMs could tune their own decoding—no more guesswork with temperature and top‑p? Enter AutoDeco: the first architecture that makes LLM decoding truly end-to-end. By adding lightweight heads, the model predicts its own context-aware temperature and top‑p at every token
LLMS
-
GPT-5 Shows Improved Routing Accuracy and Overall Performance
By
–
i feels like GPT-5’s routing accuracy and overall performance have improved a lot. Is it just me?
-

MIT’s WorldTest Benchmark Challenges Top AI Models’ World Understanding
By
–
MIT just quietly humbled every major AI lab — and almost nobody’s talking about it. They built a new benchmark called WorldTest to see if AI actually understands the world. The results are brutal.
Even the top models — Claude, Gemini 2.5 Pro, OpenAI o3 — got crushed by -
GPU Tiers Determine Hiring and AI Agent Productivity
By
–
no, you will be hired based on how many GPUs you have (and which ones as well) higher # of GPUs at higher tiers
means better LLMs
which means better Agentic tooling
and higher productivity -

DeepSeek vs Kimi 2: reasoning capabilities comparison
By
–
Bon bah y'a eu deepseek, maintenant y'a kimi 2. La différence c'est les "raisonnements"
-

Nested Optimization for Continual Learning and Long Context Processing
By
–
An exciting new approach for doing continual learning, using nested optimization for enhancing long context processing.
-

OpenAI prepares the launch of GPT-5.1
By
–
BREAKING : OpenAI is preparing to ship GPT-5.1 family with 3 new models: – GPT-5.1
– GPT-5.1 Reasoning
– GPT-5.1 Pro These traces appear in the code for Enterprise RBAC settings, and the release date is set to "November 24". What if… https://
x.com/scaling01/stat
/scaling01/status/1986886020067938749
… -
Cerebras Code Upgrades to GLM 4.6 with Record Speed
By
–
Cerebras Code just got an UPGRADE. It's now powered by GLM 4.6
— Cerebras (@cerebras) 7 novembre 2025
Pro Plans ($50):
300k ▶️ 1M TPM @ 24M Tokens/day
Max Plans ($200):
400k ▶️ 1.5M TPM @ 120M Tokens/day
Fastest GLM provider on the planet at 1000 tokens/s and at 131K context.
Get yours before we run out 👇 pic.twitter.com/WUlYPVEydTCerebras Code just got an UPGRADE. It's now powered by GLM 4.6 Pro Plans ($50):
300k 1M TPM @ 24M Tokens/day Max Plans ($200):
400k 1.5M TPM @ 120M Tokens/day Fastest GLM provider on the planet at 1000 tokens/s and at 131K context. Get yours before we run out -

ByteDance Ouro: Scaling Latent Reasoning in Language Models
By
–
New paper from ByteDance Seed: Scaling Latent Reasoning via Looped LMs This paper proposes Ouro, which reuse the same layers to think in latent space instead of dumping long chain-of-thought text 2-3x param efficiency + increased performance via iterative latent computation
-
New Compact Codex Mini Model with Increased Usage Limits
By
–
More Codex for everyone with a more compact model and increased usage limits! Try the new mini model in Codex CLI, giving 4× more usage than GPT-5-Codex:
❯ codex -m gpt-5-codex-mini
