I've long preferred Claude Code over Codex or Gemini, because it seemed much more reliable, but couldn't explain why : now Bullshit Bench by @petergostev provides compelling numbers. It measures bullshit as "when given false premises disguised in jargon, will the model go with
LLMS
-
OpenAI seeks feedback on model performance issues
By
–
Hey hey – VB from OpenAI, what's wrong? would love to know specific instances where it didn't perform as well. DMs open!
-
xAI to train three Grok Build models simultaneously this weekend
By
–
By this weekend, xAI will have three Grok Build models in training simultaneously
-

GLM-5-Turbo performance and inference speed discussion
By
–
Inference speed is the new battlefield, and GLM-5-Turbo is heavily armed. Blown away by the initial specs here. Can't wait to test the API.
-
Opus Audit Assessment of Leading AI Models Compared
By
–
I was using Opus via Cursor, did an audit with Gemini 3.1 Pro, Opus 4.6 and GPT-5.4. Then I asked Opus to give assessment of the audit quality (anonymously). And I think it 100% nailed the current state of the models: Gemini 3.1 Pro: The weakest. Looked at the screen. Found the
-
Comprehensive Roster of Current and Upcoming Large Language Models
By
–
Here is the full roster! – Llama 3 8B
– OLMo 2 7B
– DeepSeek V3
– DeepSeek R1
– Gemma 3 27B
– Mistral Small 3.1 24B
– Llama 4 Maverick
– Qwen3 235B-A22B
– Qwen3 32B
– Qwen3 8B
– Qwen3 4B
– SmolLM3 3B
– Kimi K2
– GLM-4.5 355B
– GPT-OSS 20B
– GPT-OSS 120B
– Grok 2.5 270B
– Qwen3 -

Comprehensive Visual LLM Architecture Gallery Released
By
–
Someone built the ultimate visual LLM Architecture Gallery, packing 38 models from 2024-2026 into a single hub It completely breaks down the complexity for you. Inside:
→ Annotated diagrams
→ Key design choices
→ Actual code implementations link to the gallery in ↓ -

Attention Residuals: Selective Computation in Deeper Transformers
By
–
“Attention Residuals” is now available on AlphaXiv! In standard transformer, every layer just inherits an equal sum of all earlier layers, so as models get deeper, useful computations get diluted instead of being selectively reused. The research team at @Kimi_Moonshot proposes
-
OpenAI Introduces Subagents in Codex for Parallel Codebase Management
By
–
🚨 @OpenAI just launched Subagents in Codex, and it changes how we handle massive codebases.
— Charly Wargnier (@DataChaz) 17 mars 2026
You no longer have to rely on a single agent.
You can now split complex features across multiple specialized workers running in parallel 🔥
By spinning up Codex’s new custom subagents,… pic.twitter.com/4JXFeavgXA@OpenAI just launched Subagents in Codex, and it changes how we handle massive codebases. You no longer have to rely on a single agent. You can now split complex features across multiple specialized workers running in parallel By spinning up Codex’s new custom subagents,
-

Anthropic Launches Claude Certified Architect Certification for AI Agents
By
–

There’s a new gold standard if you want to build multi-agent systems and enterprise tools. @AnthropicAI just dropped the `Claude Certified Architect` certification. It’s a rigorous 60-question, 120-minute exam that proves you can build production-grade applications. To
