We’ve post trained a model on top of Qwen that achieves Pareto optimality on accuracy-cost curves. Unlike our previous post trained models, this model has been trained to be good at search and tool calls simultaneously, allowing us to unify the tool call router and
MACHINE LEARNING
-

Reward Design Balances Correctness Preference Efficiency
By
–
Our reward design combines correctness, preference, and efficiency. Preference only counts when the answer is correct. This keeps the model from optimizing for better-sounding wrong answers.
-

Fine-tuning and On-Policy RL for Model Optimization
By
–
We first fine-tune the model to follow instructions, stay within guardrails, and keep language consistent. Then we run on‑policy RL to improve search accuracy and tool efficiency while preserving those behaviors.
-

New Research: SFT+RL Pipeline Boosts Search-Augmented AI Accuracy
By
–
We've published new research on how we post-train models for accurate search-augmented answers. Our SFT + RL pipeline improves search, citation quality, instruction following, and efficiency. With Qwen models, we match or beat GPT models on factuality at a lower cost.
-

Discovering Novel LLM Experts via Task-Capability Coevolution
By
–
Most LLM development still optimizes one model at a time. So this paper from Sakana AI proposes AC/DC (Assessment Coevolving with Diverse Capabilities), where models and tasks evolve together. The main idea is
-

Process Reward Models Grade Robot Performance Like Sports Replay
By
–
What if we could audit a robot's performance like a sports replay, grading every move, not just the final score? Researchers from Peking University, Chinese Academy of Sciences, and the Beijing Academy of AI present PRM-as-a-Judge. They use a "Process Reward Model" to watch a
-
Using temporary helper images as multimodal input for generation
By
–
Internally it appears to create temporary “helper images” (presumably using code here) and then uses those outputs as multimodal input for the final generation.
-
NVIDIA Asset Harvester Converts Autonomous Driving Video to 3D Assets
By
–
NVIDIA researchers just released Asset Harvester — an end-to-end pipeline that turns autonomous driving video into manipulable 3D object assets.
— NVIDIA AI (@NVIDIAAI) 22 avril 2026
A key building block for dynamic scene simulation in AV development. Code is open. Check it out: https://t.co/mMgk65fNDJNVIDIA researchers just released Asset Harvester — an end-to-end pipeline that turns autonomous driving video into manipulable 3D object assets. A key building block for dynamic scene simulation in AV development. Code is open. Check it out:
-

Recursive Language Models: Expert Q&A with MIT PhD Student
By
–
Ask us your questions about Recursive Language Models (RLMs) & where LMs are underutilized or inefficient. We’ll pick some to feature in an upcoming explainer w/MIT PhD student Alex Zhang (
@a1zhang
), who recently developed RLMs. For more on his work: https://
bit.ly/4vGyLxo -

DOE Under Secretary Leads AI Science Discovery Conference at Stanford
By
–
Join U.S. Department of @ENERGY Under Secretary for Science @dariogila
, together with leading scientists, engineers, and researchers, at the upcoming @StanfordHAI and Stanford Data Science conference on AI+Science: Accelerating Discovery. Register here: https://
hai.stanford.edu/events/ai-scie
nce-accelerating-discovery
…
