The benchmarks show the gap. NVLS all-reduce latency drops from 586.1µs on H200 to 313.3µs on GB200. In MoE prefill at EP=4, combine falls from 730.1µs to 438.5µs. For decode, GB200 sustains much higher throughput at high token speeds.
@perplexity_ai
-
Hardware Optimization for AI Model Prefill and Decode Phases
By
–
Prefill and decode stress hardware differently. Prefill is compute-bound, so Blackwell Tensor Cores, memory bandwidth, NVLink, and SHARP reductions help. Decode is latency/memory-bound, where GB200’s rack-scale NVLink domain opens up parallelism Hopper could not.
-

Research on Serving Qwen3 235B Models on NVIDIA GB200 Racks
By
–
We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up over Hopper for high-throughput inference on large MoE models, not just a training platform.
-
Internal Manual for Building AI Agent Skills Published
By
–
We've published our internal manual for building agent skills. Skills require a new way of thinking for developers.
-

Perplexity Computer Enables Autonomous Agentic Operations on Local Machines
By
–
Personal Computer in the Mac app allows Perplexity Computer to run continuously, autonomously, and locally. Paired with the Comet browser, it operates web-based tools without direct connectors. Local or remote, it takes agentic operations anywhere they’re needed.
-

Perplexity develops ROSE inference engine with CuTeDSL for faster GPU kernels
By
–
We’ve developed our own inference engine Runtime-Optimized Serving Engine (ROSE) to serve models ranging from embeddings to trillion-parameter LLMs. With CuTeDSL integrated into our inference engine, Perplexity can build the specialized GPU kernels faster to bring models up to
-

GPT-5.5 Launches on Perplexity and as Default Orchestration Model
By
–
GPT-5.5 is now available on Perplexity for Max subscribers. GPT-5.5 is also rolling out as the default orchestration model in Computer for both Pro and Max subscribers.
-

Moonshot Releases Kimi K2.6 Open-Weight Model
By
–
Kimi K2.6, the new state-of-the-art open-weight model from Moonshot, is now available for Pro and Max subscribers.
-
Perplexity’s Pipeline Improves Base Model Accuracy and Efficiency
By
–
This pipeline is why the same base model produces more accurate, better-cited, and more efficient answers inside Perplexity than out of the box. Read our research:
-

Reward Design Balances Correctness Preference Efficiency
By
–
Our reward design combines correctness, preference, and efficiency. Preference only counts when the answer is correct. This keeps the model from optimizing for better-sounding wrong answers.