2/4: We’ll be sharing a couple other fun and more complex demos this week where the GLM 5.2 agent conducts ablations on recent continual learning papers like SDPO. This can hopefully give you a sense of what these models can and cannot do when it comes to assisting in the
MACHINE LEARNING
-
SkyRL async RL training with autonomous research agent
By
–
1/4: A couple notes on the implementation. The async RL training itself is powered by SkyRL, with the research agent’s goal being resolving setup issues (in this case a libnuma dependency) and analyzing runs autonomously.
-
GLM 5.2: First high-performance open-weights model for auto-research
By
–
Introducing GLM 5.2 for autoresearch
— alphaXiv (@askalphaxiv) 22 juin 2026
GLM 5.2 is the first open weights model we've tried on our autoresearch pipeline that's proven capable for real research tasks.
With Fable 5's restrictions on research, having an open weights alternative is a huge win for open source
Watch… pic.twitter.com/y0kBtJzj5KIntroducing GLM 5.2 for auto-research. GLM 5.2 is the first open-weights model we tested on our auto-research pipeline that proved capable for real research tasks. With Fable 5's restrictions on research, having a
-
AI mania: Spidey feeling that models have changed
By
–
My version of AI mania is when I get a spidey feeling that the models have changed. Opus 4.8 feels very different today.
-

Strict revisit consistency trajectories test AI location memory
By
–
/7 Standard metrics aren't enough to prove an AI remembers specific locations. So, the team created strict "revisit consistency" trajectories: > Out-and-back: tests appearance stability
> Closed-loop: tests layout consistency
> Translation-rotation: tests identity preservation -

Real-time streaming inference model for interactive worlds
By
–
/5 An interactive world isn't truly interactive if it lags. This model is built specifically for real-time streaming inference. Using DMD-style distillation and an autoregressive rolling KV cache, it generates environments chunk-by-chunk from noise. When paired with asynchronous
-

DreamX-World 1.0 eliminates color drift and style mutations via long-rollout training
By
–
/6 Long-form autoregressive generation usually suffers from accumulated prediction errors, leading to color drift and style mutations. Thanks to specialized long-rollout training, DreamX-World 1.0 overcomes this limitation. It maintains stunning visual fidelity, smooth motion,
-

Announcement of new GPT-5.6, Pro, and bidirectional voice models this Thursday
By
–
It looks like we are going to have a whole range of new GPT models this Thursday: GPT-5.6, 5.6 Pro, and a new bidirectional voice model. Initial tests of the voice model have been exceptional, this is exactly what I was hoping for two years ago!
-

Largest LLM-as-Judge audit shows exact-match overstates skill
By
–
The largest LLM-as-a-Judge reliability audit yet. Researchers ran 21 judges from nine providers over roughly 541,000 judgments on MT-Bench, JudgeBench, and RewardBench. Findings: Validating a judge with exact-match agreement overstates its skill, because exact match does not
-

Prompt Reinjection fixes forgetting detailed descriptions
By
–
Why do AI image generators keep forgetting your detailed descriptions? A team from Fudan University, Alibaba, and Baidu introduces Prompt Reinjection. They found that multimodal diffusion transformers gradually lose prompt info in deeper layers. Their training-free fix: