You can also check out our training scripts here: https://
github.com/NovaSky-AI/Sky
RL/tree/main/examples/train/rlm
… Thanks to @sumanthrh and @charlie_ruan from the SkyRL team for their responsiveness in resolving issues and for providing an amazing RL library for the community! We'd also like to thank @a1zhang for his
MACHINE LEARNING
-
Open Source Reinforcement Learning Training Scripts Released
By
–
-

Reinforcing Recursive Language Models for Long-Context Tasks
By
–
Reinforcing Recursive Language Models Can a 4B model learn to recursively call itself to answer hard long-context questions? We RL fine-tuned a small model to behave as a native RLM. On evidence selection across scientific papers, our 4B RLM matches Sonnet 4.6 in quality
-

AI Competition and Governance: US vs. China and World Models
By
–
In this Q&A, @StanfordHAI Executive Director @russellwald discusses the AI competition between the U.S. and China, and the emergence of world models and their impact on AI governance. Read via @POLITICO
: https://
politico.com/newsletters/di
gital-future-daily/2026/05/08/5-questions-for-russell-wald-00911932
… -

Google DeepMind introduces AI Co-Mathematician agent for research
By
–
NEW paper from Google DeepMind. (bookmark it) AI Co-Mathematician is an agentic research workbench for mathematicians, and it just hit 48% on FrontierMath Tier 4, a new high score among AI systems evaluated. The system is an asynchronous, stateful environment that supports
-
Symbolic Learning as a Scalable Alternative to Neural Networks
By
–
Symbolic learning is not a replacement for coding agents, it's a replacement for gradient descent & NNs: a low-level, completely general, extremely scalable new learning substrate.
-

Optimizing Throughput for Large MoE Models on GB 200 Hardware
By
–
GB 200s change how one does the prefill and decode disaggregation when serving large MoEs like Qwen. We’ve published details of our stack quantifying the throughput benefits compared to serving on Hoppers.
-
NVIDIA GB200 Architecture Optimized for Large-Model Inference
By
–
This NVIDIA remains the strongest platform for large-model inference at scale. Prefill/decode disaggregation, Blackwell-native quantization, custom kernels, and rack-scale NVLink turn GB200 into faster answers lower serving cost. Read the full paper here
-
Performance Benchmarks: NVIDIA H200 vs GB200 for AI Workloads
By
–
The benchmarks show the gap. NVLS all-reduce latency drops from 586.1µs on H200 to 313.3µs on GB200. In MoE prefill at EP=4, combine falls from 730.1µs to 438.5µs. For decode, GB200 sustains much higher throughput at high token speeds.
-

Research on Serving Qwen3 235B Models on NVIDIA GB200 Racks
By
–
We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up over Hopper for high-throughput inference on large MoE models, not just a training platform.
-
Gary Marcus’s Warnings on AI Generalization Proven True
By
–
“Marcus's repeated warnings about the "wall of generalization" since 1998 have once again been proven true.”