5/5 The takeaway: If your agent relies on an LLM judge for selection accuracy, measuring code quality isn’t enough; you need a measure of the model's inductive bias toward the "fingerprint" of a gold solution. This was our blueprint.
@ai21labs
-

Upgrading LLM Reducer Prompts for Better Constraint Hierarchy
By
–
4/5 We upgraded our original 3-line “be correct” prompt → a much more detailed prompt that enforced a hierarchy of constraints for correctness, regression safety, and minimality. Basically, get the LLM-based reducer to stop being an aesthetic snob.
-
LLM Judge Rejects Functional Fix for Code Aesthetics
By
–
3/5 An example: In instance psf__requests-1724, the gold fix is 2 lines. Our agent’s functional fix was 8 lines. The LLM judge rejected the correct 8-liner as "messy" and "redundant," choosing a clean but **non-functional** fix instead. See full patch in the blog:
-
The Model Preferred Aesthetics Over Actual Function
By
–
2/5 Turns out the model wasn't remembering the solution, but it was identifying "gold-like" aesthetics like minimality & clarity. Total form over function kind of scenario.
-

Claude Opus 4.5 Judges Gold Patches Without Memorization
By
–
1/5 When we saw our Reducer (=LLM judge component in Maestro, our agentic framework, that selects the best output from parallel agent runs) consistently picking gold patches, we were sure Claude Opus 4.5 (knowledge cutoff Aug '25) had simply memorized the answers. But then we
-
LLM judges reject functional agent fixes as messy code
By
–
3/5 An example: In instance psf__requests-1724, the gold fix is 2 lines. Our agent’s functional fix was 8 lines. The LLM judge rejected the correct 8-liner as "messy" and "redundant," choosing a clean but **non-functional** fix instead. See full patch in the blog:
-
Model Aesthetics Over Function Recognition Bias
By
–
2/5 Turns out the model wasn't remembering the solution, but it was identifying "gold-like" aesthetics like minimality & clarity. Total form over function kind of scenario.
-

AI21’s Chief on Optimizing Real-World AI Agents
By
–
Catch Or Dagan, AI21's Chief Product & Strategy Officer, at #AIDev26 on April 29: "An End to Manual Tinkering: Optimizing Accuracy, Cost & Latency in Real-World Agents". We'll be at Booth 121. See you in SF! @DeepLearningAI @AndrewYNg
-
AI21 Labs: Model Scaling Hits Diminishing Returns, Orchestration Key
By
–
Scaling model size is hitting diminishing returns. The real gains are in orchestration. Our Co-Founder & Co-CEO @yshoham makes the case in a rare long-form profile by @Calcalistech today. CTech (@Calcalistech) The man trying to ground the AI boom. Prof. @yshoham, one of the world's leading artificial intelligence researchers and founder of @AI21Labs, discusses hype cycles, unreliable systems, and the long road to trustworthy artificial intelligence. calcalistech.com/ctechnews/a… — https://nitter.net/Calcalistech/status/2043339162607075708#m
-
Shoham Criticizes Utopia and Apocalypse, Advocates for Reliable AI
By
–
"Both the utopia and apocalypse scenarios ignore AI's limitations." Prof. Yoav Shoham, one of the world's leading artificial intelligence researchers and founder of @AI21Labs, discusses hype cycles, unreliable systems, and the long road to trustworthy AI.
calcalistech.com/ctechnews/a… [Translated from EN to English]
