
It also sets a new cost-performance frontier. On DSQA it scores 0.871, ahead of Anthropic's 0.815, at nearly half the cost per task. On WideSearch it leads on score while running cheaper.

By
–

It also sets a new cost-performance frontier. On DSQA it scores 0.871, ahead of Anthropic's 0.815, at nearly half the cost per task. On WideSearch it leads on score while running cheaper.

By
–
We tested Search as Code on deep research (DSQA, BrowseComp, HLE) and wide research benchmarks (WideSearch, WANDR). It matches or beats every competing system across all five.

By
–
The traditional tool-calling approach suffers from high latency, manual control flow, and context pollution. With Search as Code, the model composes search primitives: fanning out queries asynchronously, deduping, filtering, joining, and ranking before results hits its context.

By
–
Most LLM search still works by sampling more rollouts or extending one path at a time. This paper's bidirectional evolutionary search does it in a smarter way. It breaks the task backward into smaller

By
–
How to build AI agents from scratch (9 steps): 1. Purpose & scope 2. I/O schemas 3. System instructions 4. Reasoning + tools 5. Multi-agent orchestration 6. Memory & context 7. Multimodal 8. Structured outputs 9. UI / API Ship agents that do work, not just talk.
By
–
This pod was an incredible gift to the community:
— swyx (@swyx) 1 juin 2026
not only our first pod about @xAI, but Ethan really indulged on all our questions on how to train a SOTA Videogen world model, including specific areas (consistent extending/editing, voice) that Grok @Imagine is *still* SOTA,… https://t.co/purQe9Ja81 pic.twitter.com/Sl4AqAt7RA
This pod was an incredible gift to the community: not only our first pod about @xAI
, but Ethan really indulged on all our questions on how to train a SOTA Videogen world model, including specific areas (consistent extending/editing, voice) that Grok @Imagine is *still* SOTA,
By
–
longmemeval experiment arch: 1) deterministic ingestion/extraction (85.6% accuracy, 86.2% retrieval)
2) semantic ingestion/extraction (84.8% accuracy, 94.9% retention)
3) semantic ingestion/deterministic extraction (87.6% accuracy, 90.0% retrieval) this was staggered, not
By
–
“Test-time compute” is such a dumb name for “compute”. So many dumb names for things in AI.
By
–
Yeah, but opensource and weights and 1m context. So I’d say it’s a win
By
–
Grok Imagine’s Agent Moment: Cosmos, xAI, World Models, Generative UI, & the Codex Phase for Video! https://
latent.space/p/video-agents @EthanHe_42
, former @xai world model lead and @nvidia Cosmos researcher, explains why AI video may follow the same path as coding agents, how Grok