The models are getting smarter and more confident at the same pace. Your verification system is now more valuable than your prompting system.
LLMS
-

GPT-5.5: Smartest model but most confidently wrong benchmark results reveal flaw
By
–
GPT-5.5 is the smartest model ever tested. It's also the most confidently wrong. That's not an opinion. That's what the benchmarks say when you read both columns. Artificial Analysis runs AA-Omniscience, a benchmark designed to penalize models that guess instead of saying "I
-
MIT PhD Student Calls AI Models Mismanaged Geniuses With Broader Potential
By
–
MIT PhD student Alex Zhang (@a1zhang) explains how AI models are "mismanaged geniuses" that could take on a much wider range of tasks.
— MIT CSAIL (@MIT_CSAIL) 30 avril 2026
Full video: https://t.co/8L9lVGtzF1 pic.twitter.com/G38iDOgS1DMIT PhD student Alex Zhang (
@a1zhang
) explains how AI models are "mismanaged geniuses" that could take on a much wider range of tasks. Full video: https://
tinyurl.com/bddd5vdx -
Frontier vs Capable AI Models: A Quick Comparison Guide
By
–
Absolute frontier Claude Opus 4.x / Sonnet 4.x
GPT-5 / GPT-4.1 family
Gemini 2.5 / 3 Pro
DeepSeek R1 (reasoning) Mistral 3 large Very capable
Often cheaper / efficient
But not SOTA in intelligence -
AI Companies and Projects Mentioned in Article
By
–
France-based SquareMind just raised $18 M for Swan, an AI robotic platform that automates full-body skin imaging to help detect skin cancer.
— The Rundown AI (@TheRundownAI) 30 avril 2026
A cool look at the tech-enabled future of dermatology: pic.twitter.com/8RKjJBOUOjFrance-based SquareMind just raised $18 M for Swan, an AI robotic platform that automates full-body skin imaging to help detect skin cancer. A cool look at the tech-enabled future of dermatology:
-
alphaXiv Partners with OpenRouter for Direct Model Access
By
–
alphaXiv 🤝 OpenRouter
— alphaXiv (@askalphaxiv) 30 avril 2026
Excited to announce our partnership with @OpenRouter
You can now hover over any model name on any paper, and you’ll be directly connected to OpenRouter.
In the pop up window, you’ll be able to see the provider name, model name, description, and its top… pic.twitter.com/cwOC1Cd7fHalphaXiv OpenRouter Excited to announce our partnership with @OpenRouter You can now hover over any model name on any paper, and you’ll be directly connected to OpenRouter. In the pop up window, you’ll be able to see the provider name, model name, description, and its top
-
Frontier Model APIs vs Native Apps: A Growing Capability Gap
By
–
Increasingly, I think, we will see a gap between what you can do with frontier model APIs & what you can do with the native apps from the frontier labs (Codex, Claude Code). Models developed and trained with their native harnesses in mind have more capabilities in their harnesses
-
Codex and GPT-5.5 Updates Incoming to Maximize Capabilities
By
–
love to see it, lots more updates enroute to help you maximise what you can do with Codex & GPT 5.5 if you have any feedback please do send it my way
-
LLM Refusals and Context Window Name-Dropping Skepticism
By
–
"Why are you so impressed? We know LLMs can coherently namecheck things alluded to in a context window." The refusal is 1,400 words long.
-

AgentTrove: New Agentic Dataset with 1.7M Samples Released
By
–
AgentTrove: new agentic dataset with 1.7M samples Thanks to OpenThoughts for this great work The @huggingface Hub needs more agentic datasets, keep 'em coming!