Watch every session from Interrupt, the agent conference by LangChain.
LLMS
-
PerceptionDLM: Parallel Region Perception with Multimodal Language Models
By
–
PerceptionDLM
— AK (@_akhaliq) 22 juin 2026
Parallel Region Perception with Multimodal Diffusion Language Models pic.twitter.com/0vZdGaAPoyPerceptionDLM Parallel Region Perception with Multimodal Diffusion Language Models
-
Play with SkyRL via OpenResearch.sh or autoarxiv with GLM 5.2
By
–
4/4: You can play with this yourself! Visit http://
OpenResearch.sh (
http://
openresearch.sh) or change ‘arxiv’ to ‘autoarxiv’ on the official SkyRL paper https://
autoarxiv.org/abs/2511.16108 and use the GLM 5.2 model to iterate on the repo! -
GLM 5.2 lacks image understanding; uses numpy for WandB charts
By
–
3/4: One limitation worth noting: GLM 5.2 has no image understanding. While Opus and Fable can consistently identify trends in WandB charts, GLM resorts to writing numpy code to smooth and clean the raw WandB numbers before analyzing. For simpler runs like this example this is
-
GLM 5.2 agent ablation demos on continual learning papers
By
–
2/4: We’ll be sharing a couple other fun and more complex demos this week where the GLM 5.2 agent conducts ablations on recent continual learning papers like SDPO. This can hopefully give you a sense of what these models can and cannot do when it comes to assisting in the
-
GLM 5.2: First high-performance open-weights model for auto-research
By
–
Introducing GLM 5.2 for autoresearch
— alphaXiv (@askalphaxiv) 22 juin 2026
GLM 5.2 is the first open weights model we've tried on our autoresearch pipeline that's proven capable for real research tasks.
With Fable 5's restrictions on research, having an open weights alternative is a huge win for open source
Watch… pic.twitter.com/y0kBtJzj5KIntroducing GLM 5.2 for auto-research. GLM 5.2 is the first open-weights model we tested on our auto-research pipeline that proved capable for real research tasks. With Fable 5's restrictions on research, having a
-
AI mania: Spidey feeling that models have changed
By
–
My version of AI mania is when I get a spidey feeling that the models have changed. Opus 4.8 feels very different today.
-

Announcement of new GPT-5.6, Pro, and bidirectional voice models this Thursday
By
–
It looks like we are going to have a whole range of new GPT models this Thursday: GPT-5.6, 5.6 Pro, and a new bidirectional voice model. Initial tests of the voice model have been exceptional, this is exactly what I was hoping for two years ago!
-

Largest LLM-as-Judge audit shows exact-match overstates skill
By
–
The largest LLM-as-a-Judge reliability audit yet. Researchers ran 21 judges from nine providers over roughly 541,000 judgments on MT-Bench, JudgeBench, and RewardBench. Findings: Validating a judge with exact-match agreement overstates its skill, because exact match does not
-
Sakana Fugu matches Fable, Mythos with single API, Japan in big league
By
–
I don't think anyone saw this coming.
— Charly Wargnier (@DataChaz) 22 juin 2026
Sakana Fugu matches Fable and Mythos through a single model API 🤯
Japan is officially in the big league. https://t.co/DIdwrwYFdm pic.twitter.com/I9qmc9xS3II don't think anyone saw this coming. Sakana Fugu matches Fable and Mythos through a single model API Japan is officially in the big league.