We've published new research on how we post-train models for accurate search-augmented answers. Our SFT + RL pipeline improves search, citation quality, instruction following, and efficiency. With Qwen models, we match or beat GPT models on factuality at a lower cost.
RESEARCH
-

Discovering Novel LLM Experts via Task-Capability Coevolution
By
–
Most LLM development still optimizes one model at a time. So this paper from Sakana AI proposes AC/DC (Assessment Coevolving with Diverse Capabilities), where models and tasks evolve together. The main idea is
-

Process Reward Models Grade Robot Performance Like Sports Replay
By
–
What if we could audit a robot's performance like a sports replay, grading every move, not just the final score? Researchers from Peking University, Chinese Academy of Sciences, and the Beijing Academy of AI present PRM-as-a-Judge. They use a "Process Reward Model" to watch a
-
MC0001 Conference on Machine Consciousness Philosophy May 2026
By
–
If you are interested in the philosophy of consciousness and machine consciousness research, you should come! MC0001 Conference (May 29-31, Lighthaven, Berkeley)
-
Using temporary helper images as multimodal input for generation
By
–
Internally it appears to create temporary “helper images” (presumably using code here) and then uses those outputs as multimodal input for the final generation.
-

AI-Generated Fraudulent Paper: Academic Integrity Crisis Exposed
By
–
This is a FRAUDULENT paper, AI-generated. My name was used as an author and I had nothing to do with it, never saw it until today https://
e-pubmed.co.uk/journals/digit
al-health-implementation/articles/implementation-science-ai-digital-health-systematic-review/
…
The "Editors" Angelo Rossi Mori, David Mensah, and Zarnie Khadjesari should be reported. -
Economic hopes and worries from 81,000 people on AI
By
–
Last month, we published our look into what 81,000 people told us they want from AI. In new research, we’ve investigated the economic hopes and worries referenced in their responses. Read more: https://
anthropic.com/research/81k-e
conomics
… -
Evaluating AGI: Beyond Mimicry to True Learning Capability
By
–
Judging AGI by how well it can mimic us is a category error, because mimicry isn't intelligence and isn't general. We should judge AGI by how well it learns to do things we didn't teach it (including things we don't know how to do ourselves).
-

AI Systems Can Now Generate Credible Papyrus Documents
By
–
on peut même faire des papyrus crédibles
-

Building Research Excellence: Global Insights Report
By
–
What shapes #research excellence? Elsevier's latest report, "How research environments and practices foster excellence", shows that excellence is intentionally built.
Based on insights from 100+ #researchers, leaders, and funders across Brazil, China, Germany, and India, it
