How could anyone trust this guy, his work, benchmarks, or any kind of output / product in Local AI after this is beyond me at this point “Any good benchmark will be trained on since that’s what people want.” I rest my case
RESEARCH
-

Humans as dice: LLMs lack human variation
By
–
The Matrix idea of keeping humans as batteries is obviously weird… we would be more useful as dice. LLMs default to very similar kinds of arguments & structure, and even different LLMs seem to collapse to similar concepts. Humans provide a lot more variation in their own work.
-

Llama 3 8B and 405B trained with NVFP4, up to 1.73x faster
By
–
We trained Llama 3 8B and 405B with NVFP4 precision on the NVIDIA Blackwell platform. Here's what we found: 1.31–1.73× faster than FP8, with zero accuracy loss.
-
AI Consciousness Discussed with Geoff Nielson on Podcast
By
–
Could AI already be conscious? Geoff Nielson has been talking to me on the Digital Disruption podcast https://
youtube.com/watch?v=VVc0hi
A9O7c
… -
Honest evals without self-gold medals to motivate team under public pressure
By
–
i was very keen on publishing evals where we do not award ourselves gold medals, as an intellectual honesty thing and motivator for the team to hillclimb in ways we care about with public pressure
-

OpenAI enters third phase as economy reshapes around AI
By
–

OpenAI is "entering the third phase. The economy is beginning to reshape around AI." – The first phase of OpenAI was about doing research toward AGI
– The second phase began when the research became relevant to the real world and OpenAI became a product company Their goal for -

METR report: most SWEBench results unmergeable; FrontierCode unsolved by frontier models
By
–

It's finally out!!! @METR_Evals found that more than half of SWEBench results is unmergeable slop. FrontierCode represents over 1000+ hours of maintainer validated software engineering work most frontier models cannot yet solve, much less solve with high quality. Cog had IOI
-

Agentopia: Long-Running Life Simulation for LLM Agent Societies
By
–
// Life Simulation in Agent Societies // One of the more ambitious agent-society testbeds to land this month, and it arrives as a 79-page release. Agentopia drops many LLM agents into a long-running world where they live, interact, and learn over extended horizons. The goal is
-
Snorkel AI Reading Group on JudgmentBench with Russell Yang
By
–
Join us for late-afternoon boba and research. RSVP: https://t.co/lMUBYMGPZx🧋
— Snorkel AI (@SnorkelAI) 8 juin 2026
Next up in the Snorkel AI Reading Group: Russell Yang (@StanfordLaw) on “JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment,” stemming from a collaboration with @harvey… pic.twitter.com/bexEc52RVXJoin us for late-afternoon boba and research. RSVP: https://
luma.com/qwsxqfa2 Next up in the Snorkel AI Reading Group: Russell Yang (
@StanfordLaw
) on “JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment,” stemming from a collaboration with @harvey -
Snorkel AI Reading Group: Russell Yang on JudgmentBench
By
–
Join us for late-afternoon boba and research. RSVP: https://t.co/lMUBYMGPZx🧋
— Snorkel AI (@SnorkelAI) 8 juin 2026
Next up in the Snorkel AI Reading Group: Russell Yang (@StanfordLaw) on “JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment,” stemming from a collaboration with @harvey… pic.twitter.com/bexEc52RVXJoin us for late-afternoon boba and research. RSVP: https://
luma.com/qwsxqfa2 Next up in the Snorkel AI Reading Group: Russell Yang (
@StanfordLaw
) on “JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment,” stemming from a collaboration with @harvey
