We’ve post trained a model on top of Qwen that achieves Pareto optimality on accuracy-cost curves. Unlike our previous post trained models, this model has been trained to be good at search and tool calls simultaneously, allowing us to unify the tool call router and
LLMS
-
Perplexity’s Pipeline Improves Base Model Accuracy and Efficiency
By
–
This pipeline is why the same base model produces more accurate, better-cited, and more efficient answers inside Perplexity than out of the box. Read our research:
-

Reward Design Balances Correctness Preference Efficiency
By
–
Our reward design combines correctness, preference, and efficiency. Preference only counts when the answer is correct. This keeps the model from optimizing for better-sounding wrong answers.
-

Fine-tuning and On-Policy RL for Model Optimization
By
–
We first fine-tune the model to follow instructions, stay within guardrails, and keep language consistent. Then we run on‑policy RL to improve search accuracy and tool efficiency while preserving those behaviors.
-

New Research: SFT+RL Pipeline Boosts Search-Augmented AI Accuracy
By
–
We've published new research on how we post-train models for accurate search-augmented answers. Our SFT + RL pipeline improves search, citation quality, instruction following, and efficiency. With Qwen models, we match or beat GPT models on factuality at a lower cost.
-
Using LLMs to Generate Safe Runnable Code
By
–
I want an LLM to generate code for me which I can then safely run somewhere without worrying about it breaking anything
-

Discovering Novel LLM Experts via Task-Capability Coevolution
By
–
Most LLM development still optimizes one model at a time. So this paper from Sakana AI proposes AC/DC (Assessment Coevolving with Diverse Capabilities), where models and tasks evolve together. The main idea is
-
GPT-Image-2 Prompt Generation and DALL-E 3 Comparison
By
–
Presumably GPT-imagegen-2 (aka ChatGPT Images 2.0 aka gpt-image-2) works as a tool which the models generate prompts for? I wish we could see those prompts, like back in the DALL-E 3 days https://
simonwillison.net/2023/Oct/26/ad
d-a-walrus/
… -
AI Acts as World-Class Research Assistant for Arcane Trivia
By
–
… this is your standard citation.” It’s really weird to have a spiky, tireless research gopher who is literally world class on what even specialists would usually describe as arcane trivia.
