I am excited to share that I have started a new adventure at @MistralAI, a leading frontier lab, where I am working on pushing further the agentic reasoning capabilities of LLMs.
→ View original post on X — @arthurmensch, 2026-03-05 14:09 UTC
By
–
I am excited to share that I have started a new adventure at @MistralAI, a leading frontier lab, where I am working on pushing further the agentic reasoning capabilities of LLMs.
→ View original post on X — @arthurmensch, 2026-03-05 14:09 UTC
By
–
AI Engineering are going to love this!
— Akshay 🚀 (@akshay_pachaar) 5 mars 2026
ART (Agent Reinforcement Trainer) is an open-source framework for training agents with GRPO + RULER (an automatic reward system).
No need to hand-craft reward functions.
GitHub: https://t.co/O4dIDm6cqv https://t.co/OHeohQTOd4 pic.twitter.com/LLhoe6nUq7
AI Engineering are going to love this! ART (Agent Reinforcement Trainer) is an open-source framework for training agents with GRPO + RULER (an automatic reward system). No need to hand-craft reward functions. GitHub: http://
github.com/OpenPipe/ART

By
–
RL isn’t the only way to fine-tune LLMs In our most recent AI4Science talk “Evolution Strategies at Scale”, Xin Qiu (
@realVsonicV
), Principal Research Scientist and Senior Director at Cognizant AI Lab, walked through the first research showing that Evolution Strategies can be
By
–
Good one! Intermediate Layer Distillation is another technique that I’ve read about. Instead of only matching the teacher’s final output distribution, you also align intermediate hidden states, attention patterns, or feature maps across corresponding layers. The intuition

By
–



How do you train a trillion-parameter AI model while dramatically improving efficiency? http://
YuanLab.ai @YuanAI_Lab presents Yuan3.0 Ultra to tackle exactly that. They introduced Layer-Adaptive Expert Pruning (LAEP) for pre-training — a system that monitors how much

By
–
“Speculative Speculative Decoding” Even though speculative decoding speed up LLM inference, it still has a hidden “stop-and-wait” step, where the draft model can’t start the next guess until the big model finishes verifying the last one. This paper fixes that by making drafting
By
–
Congrats @willjhliang, @JasonMa2020 @dineshjayaraman and collaborators on this cool combination of VLMs, keypoint detectors, and GOFE trajectory warping to self-supervise and reset 1000s of diverse pick-and-place demonstrations that are then used to train a VLA policy. https://t.co/OzEp63lx99
— Ken Goldberg (@Ken_Goldberg) 5 mars 2026
Congrats @willjhliang, @JasonMa2020 @dineshjayaraman and collaborators on this cool combination of VLMs, keypoint detectors, and GOFE trajectory warping to self-supervise and reset 1000s of diverse pick-and-place demonstrations that are then used to train a VLA policy. Will Liang (@willjhliang) Introducing Tether 🪢, a fun little idea to scale data by having our robot “play” in the real world for over 24 hours, throughout the day and overnight—improving policies from zero to mastery with minimal supervision! But play is messy, with out-of-distribution scenarios that are hard to anticipate. To perform autonomous functional play in the real world, from just a handful of demos, we propose a highly robust few-shot imitation method that warps demo trajectories using visual correspondences. Then, continuously running it within a multi-task VLM-guided cycle, we generate a data stream that produces 1000+ expert-level demos. This generated data is finally funneled downstream to train imitation learning policies, which improve from zero to near-perfect success rates. We’ll be presenting Tether at #ICLR2026 in just a few weeks! But before that, deep dive with me… 🧵 — https://nitter.net/willjhliang/status/2029238456766087386#m
→ View original post on X — @ken_goldberg, 2026-03-05 06:35 UTC

By
–


BREAKING : OpenAI has started testing a new model named “Galapagos” on Arena which potentially could be a GPT-5.4 low effort version. “Sooner than you think”

By
–
most people building AI agents obsess over how they WRITE memories
turns out that's basically irrelevant new research analyzed 9 different memory systems across 1,540 questions the finding?
retrieval method drives 20-point accuracy swings
write strategy? 3–8 points max raw
By
–
The @IEEE Transactions on Robot Learning (T-RL) will launch on March 30! Co-EiCs: Todd Murphey and Vincent Vanhoucke ieee-ras.org/publications/t-…
→ View original post on X — @ken_goldberg, 2026-03-04 19:51 UTC