4/5 We upgraded our original 3-line “be correct” prompt → a much more detailed prompt that enforced a hierarchy of constraints for correctness, regression safety, and minimality. Basically, get the LLM-based reducer to stop being an aesthetic snob.
RESEARCH
-

LLM Judge Bias: Beyond Code Quality Metrics
By
–
5/5 The takeaway: If your agent relies on an LLM judge for selection accuracy, measuring code quality isn’t enough; you need a measure of the model's inductive bias toward the "fingerprint" of a gold solution. This was our blueprint.
-
The Model Preferred Aesthetics Over Actual Function
By
–
2/5 Turns out the model wasn't remembering the solution, but it was identifying "gold-like" aesthetics like minimality & clarity. Total form over function kind of scenario.
-

Claude Opus 4.5 Judges Gold Patches Without Memorization
By
–
1/5 When we saw our Reducer (=LLM judge component in Maestro, our agentic framework, that selects the best output from parallel agent runs) consistently picking gold patches, we were sure Claude Opus 4.5 (knowledge cutoff Aug '25) had simply memorized the answers. But then we
-

Virtual Embryo and Collaborative Intelligence Research Moonshots
By
–
Learn more about their research plans: The Virtual Embryo Moonshot: https://
laude.org/moonshots/one/
virtual-embryo-moonshot
… Collaborative Intelligence for the Future of Work: https://
laude.org/moonshots/one/
collaborative-intelligence-future-of-work
… (3/3) -

Stanford Teams Win AI Research Grants and Recognition Awards
By
–
Out of 125 proposals, 4 @Stanford teams, which included senior fellows @erikbryn
, @YejinChoinka & @guestrin
; & faculty affiliates @james_y_zou
, @tatsu_hashimoto
, @Diyi_Yang
, @jure & @tengyuma were selected as: 2 seed-grant winners 1 runner-up 1 honorable mention (2/3) -

Stanford HAI Members Win Moonshots AI Research Competition Award
By
–
Congrats to members of the @StanfordHAI community recognized in the #Moonshots research competition launched by @LaudeInstitute
! Moonshots asked a simple question: how should AI actually be used to solve humanity’s hardest problems? (1/3) -
Model Aesthetics Over Function Recognition Bias
By
–
2/5 Turns out the model wasn't remembering the solution, but it was identifying "gold-like" aesthetics like minimality & clarity. Total form over function kind of scenario.
-

Being Right Too Soon: Social Acceptance of Innovation
By
–
So true. “Being right too soon is socially unacceptable”
-

Interesting AI Safety Approach Gains Attention
By
–
Definitely one of the more interesting approaches to AI safety I've seen recently
