Wow, fixing one simple parameter could stop your AI training from collapsing! UCL, Shanghai Jiao Tong University, and HKUST (Guangzhou) present HölderPO. Instead of summing token probabilities in a fixed way, HölderPO uses a flexible averaging trick controlled by a single “p”
RESEARCH
-

Peking University Researchers Unveil SEAlign for AI Code Agents
By
–
Why can't top code models handle real-world software engineering? Researchers from Peking University unveil SEAlign — a new alignment framework that trains code agents on actual software workflows. Instead of just solving coding puzzles, it uses Monte Carlo Tree Search to
-

AI capability growth past exponential takeoff per METR and UK AISA
By
–


Everyone has seen the @waitbutwhy cartoon of AI capability growth with a "you are here" indicator just before the exponential really starts, but the independent assessments of both METR and the UK's AISA do seem to show that we are past that point now (until we hit a slowdown?)
-

ProgramBench: A New Benchmark for Evaluating AI Agents in Software Development
By
–
Can AI build an entire software project from scratch, not just fix one bug? Researchers at Meta FAIR, Stanford, and Harvard introduce ProgramBench. This benchmark tests if language-model agents can take a program’s documentation and build a full codebase that behaves
-

Long-Horizon Reasoning in LLMs: An alphaXiv AI4Science Talk
By
–
If models can think for 100,000 tokens, why do they still lose the plot? Come join us for this AI4Science on alphaXiv talk: Long-Horizon Reasoning in LLMs. In this session, Sumeet Motwani (
@sumeetrm
) and Charles London (
@CharlieLondon02
) will share recent work on both training -

How Agentic LLMs Are Transforming Scientific Research and Mentorship
By
–
Why LLMs Aren't Scientists Yet? In our latest AI4Science talk, Prof. Dhruv Kumar (
@gargdhruv36
) and Dhruv Trehan (
@dhruvtrehan9
) from @lossfunk discussed how agentic LLM systems can support science in a whole new way, from generating research ideas to mentoring young researchers -
OpenAI GPT-5.6 close, question about Anthropic Sonnet 4.8
By
–
Seriously, OpenAI is on a run. GPT-5.6 very close. Where is even sonnet 4.8 @AnthropicAI ?
-
Debating the definition of world models in AI systems
By
–
I do not think there is a useful defensible definition of “no world model” which survives a model being able to position a character in three space then depict that character in a mirror in the same three space. Many people, who unquestionably have world models, would struggle.
-

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives
By
–
CausalCine Real-Time Autoregressive Generation for Multi-Shot Narratives
-

AI for Organizations Grand Challenge Explores Future of Work
By
–
200+ academic teams from 156 universities just competed to answer one question: How will AI actually change the way we work together? @StanfordHAI and @GoogleDeepMind announce winners of the AI for Organizations Grand Challenge: https://
hai.stanford.edu/news/researche
rs-worldwide-compete-to-shape-the-future-of-ai-in-organizations
…
