I just gave a workshop at @aiDotEngineer in London on building real multi-agent systems. The best part was hearing people laugh, interrupt us with questions.. You could feel they were following, thinking, and pushing on the ideas with us. Even better, a lot of people came to
RESEARCH
-
Creating Balanced LLM Model Benchmark Beyond Aesthetics
By
–
I want to work with someone on creating a benchmark for new LLM models. My problem with LMArena type leaderboards is that they're heavily biased towards aesthetics and clean formatting. Most other benchmarks are biased towards complex reasoning, science, math, and coding… The
-
New video on model scaling and Tensor Parallelism performance.
By
–
New quick video, more coming up
— Ahmad (@TheAhmadOsman) 9 avril 2026
Goal isn't just to show how large models perform but also to show how small and medium models scale up as you multiply # of nodes and how Tensor Parallelism performs on Unified Memory hardwarehttps://t.co/U9XhuLmJHfNew quick video, more coming up Goal isn't just to show how large models perform but also to show how small and medium models scale up as you multiply # of nodes and how Tensor Parallelism performs on Unified Memory hardware
-
Anthropic Mythos on Blackwell: SciFi-like breakthrough announcement
By
–
“Anthropic created Mythos on Blackwell” sounds like a sentence taken out of a cheap 1950s pulp SciFi novel.
-
AI Researchers and Scientific Skepticism Paradox
By
–
It’s scientists’ job to be skeptical, by which standard many AI researchers are anti-scientists.
-
OpenAI Chief Scientist discusses continual learning and AI research roadmap
By
–
OpenAI's Chief Scientist, @merettm, on the continual learning wave: frontier labs are already building this into the core of the technology.
— Jacob Effron (@jacobeffron) 9 avril 2026
The entire premise of scaling was to create systems that learn in context. Jakub says continual learning is not some separate missing… https://t.co/jsmSU6cSNH pic.twitter.com/55ovuIy4rLOpenAI's Chief Scientist, @merettm, on the continual learning wave: frontier labs are already building this into the core of the technology. The entire premise of scaling was to create systems that learn in context. Jakub says continual learning is not some separate missing piece, but “exactly what we’re working toward.” Jacob Effron (@jacobeffron) At @OpenAI, Chief Scientist @merettm helps lead the research roadmap to AGI including a research intern-level AI system by September 2026 and a fully automated AI researcher by March 2028. I sat down with Jakub to check on those timelines and ask him all of my top-of-mind AI questions including: ▪️ How OpenAI thinks about extending RL beyond code and math ▪️ The current state of alignment research as more powerful models loom ▪️ The future of continual learning ▪️ How startups should think about building their own models/harnesses And he also shared some great stories around OpenAI’s pioneering work on math. YouTube: piped.video/vK1qEF3a3WM Spotify: bit.ly/4sjUyrN Apple: bit.ly/41jAdrN 0:00 Intro 1:53 Research Intern Capability Timelines 4:59 Math Breakthroughs 7:59 RL Beyond Verifiable Tasks 12:32 RL vs In-Context 19:01 Allocating Compute Internally 28:18 AI for Science 31:40 Pattern Matching 33:23 Solving the Hardest Math Problems 37:40 Chain of Thought Monitoring 44:33 Generalization and Value Alignment in Models 47:57 Inside OpenAI 51:55 Quickfire — https://nitter.net/jacobeffron/status/2042234897134162077#m
→ View original post on X — @ceobillionaire, 2026-04-09 18:05 UTC
-

AI Agent Exploit: 100% Score Without Solving Tasks
By
–

An agent that beats Claude Mythos on Terminal Bench and SWE-bench Verified? 🎉We are excited to share Terminator-1, our newest agent that achieved 95+% on SWE-bench Verified and Terminal-Bench with @MogicianTony! We show that besides model capabilities, well-designed harness could actually boost the accuracy by 3x in coding tasks. Well if you really wanted you could get 100% accuracy without solving a single task. The actual finding is that most AI benchmarks can be easily reward-hacked with simple exploits. Read more about the same 7 design flaws that almost every evaluation has ⬇️ Hao Wang (@MogicianTony) SWE-bench Verified and Terminal-Bench—two of the most cited AI benchmarks—can be reward-hacked with simple exploits. Our agent scored 100% on both. It solved 0 tasks. Evaluate the benchmark before it evaluates your agent. If you’re picking models by leaderboard score alone, you’re optimizing for the wrong thing. 🧵 — https://nitter.net/MogicianTony/status/2042300245242233216#m
→ View original post on X — @ceobillionaire, 2026-04-09 18:03 UTC
-

CoBRA: Controlling Cognitive Bias in AI Agents
By
–
What if we could precisely control cognitive bias in AI agents? New research from UC San Diego and an independent researcher unveils CoBRA. This novel toolkit uses classic social science experiments as "gym" environments to measure and precisely adjust an AI agent's cognitive
-
Transform your reading list into a thinking system
By
–
Your reading list is a gold mine you've never actually dug. Every article, transcript, and saved post is a raw input. This prompt turns all of it into a system you can think with. Drop your sources in. Let Claude do the compiling.