Hölder Policy Optimisation Paper: https://
arxiv.org/abs/2605.12058
Code: https://
github.com/YihangChen9/Ho
lderPO
…
LLMS
-

Hölder Policy Optimisation paper and code
By
–
-

HölderPO: single-parameter fix to prevent AI training collapse
By
–
Wow, fixing one simple parameter could stop your AI training from collapsing! UCL, Shanghai Jiao Tong University, and HKUST (Guangzhou) present HölderPO. Instead of summing token probabilities in a fixed way, HölderPO uses a flexible averaging trick controlled by a single “p”
-

Peking University Researchers Unveil SEAlign for AI Code Agents
By
–
Why can't top code models handle real-world software engineering? Researchers from Peking University unveil SEAlign — a new alignment framework that trains code agents on actual software workflows. Instead of just solving coding puzzles, it uses Monte Carlo Tree Search to
-
Operational Usage of AI Agent Teams
By
–
Agent teams count as interactive usage, so draw from your sub
-
LangChain introduces SmithDB and LangSmith Engine for agent observability
By
–
which was your favorite launch? SmithDB (database purpose built for agent trace data): https://
langchain.com/blog/introduci
ng-smithdb
… LangSmith Engine (agent for improving your agents based on trace data): https://
langchain.com/blog/introduci
ng-langsmith-engine
… -

ProgramBench: A New Benchmark for Evaluating AI Agents in Software Development
By
–
Can AI build an entire software project from scratch, not just fix one bug? Researchers at Meta FAIR, Stanford, and Harvard introduce ProgramBench. This benchmark tests if language-model agents can take a program’s documentation and build a full codebase that behaves
-

Long-Horizon Reasoning in LLMs: An alphaXiv AI4Science Talk
By
–
If models can think for 100,000 tokens, why do they still lose the plot? Come join us for this AI4Science on alphaXiv talk: Long-Horizon Reasoning in LLMs. In this session, Sumeet Motwani (
@sumeetrm
) and Charles London (
@CharlieLondon02
) will share recent work on both training -
Fastokens addresses tokenization bottlenecks in AI inference pipelines
By
–
Thanks for building with us @CrusoeAI As context windows explode, tokenization is becoming a major hidden bottleneck in inference pipelines. fastokens is open source, already integrated with Dynamo & @lmsysorg
, and designed for the next generation of 100K-token agent -

How Agentic LLMs Are Transforming Scientific Research and Mentorship
By
–
Why LLMs Aren't Scientists Yet? In our latest AI4Science talk, Prof. Dhruv Kumar (
@gargdhruv36
) and Dhruv Trehan (
@dhruvtrehan9
) from @lossfunk discussed how agentic LLM systems can support science in a whole new way, from generating research ideas to mentoring young researchers