“Self-Distilled Agentic RL” Agent RL learns from sparse trajectory rewards, while self-distillation gives dense token guidance. But in multi-turn agents, naive distillation can break because privileged teacher signals get noisy as trajectories drift. The key idea of this paper
MACHINE LEARNING
-
AI Agents Hermes and OpenClaw Compared on GitHub History Analysis
By
–
Atomic Bot put Hermes and OpenClaw head-to-head on the exact same task, running the same model (Qwen 3.6 35B) with the same goal: analyzing GitHub history, mapping growth spikes, and shipping a live dashboard in the browser.
— 🚨 AI News | TestingCatalog (@testingcatalog) 15 mai 2026
Key metrics to watch for 👀
> Time to complete the… https://t.co/VReoAL9Taz pic.twitter.com/GgYHjaO30CAtomic Bot put Hermes and OpenClaw head-to-head on the exact same task, running the same model (Qwen 3.6 35B) with the same goal: analyzing GitHub history, mapping growth spikes, and shipping a live dashboard in the browser. Key metrics to watch for > Time to complete the
-
AI Agents Learn and Improve Daily, Aggregating AI News
By
–
My agents read the entire AI community here and make this: https://
alignednews.com/ai My agents get smarter every day because of this -
Codex generalizes problems, limiting non-coding work unnecessarily
By
–
Another aspect of this is that Codex, like a good programmer, wants to generalize problems. It has a tendency to write a repeatable code base that generates the required output. But for a lot of non-coding work, this is unnecessary and often limiting, since it anchors the work.
-

LLM Failure Modes in Learning Negations During Fine-tuning
By
–
“Negation Neglect: When models fail to learn negations in training” LLMs can understand a disclaimer in-context, but often fail to learn it during finetuning. So when training on documents saying a claim is false can still implant the claim as true. Qwen3.5 belief in
-
Runway’s Agent mode builds complex stories from short text
By
–
Fine, you all want to code like this I guess.
— Ethan Mollick (@emollick) 15 mai 2026
(Runway's new Agent mode is quite impressive, doing fairly complex story building from just a short text description of what you want. Not error free obviously, but this was pretty great for a one-shot attempt) https://t.co/vPNUwJI0gA pic.twitter.com/dO8lwbpgjKFine, you all want to code like this I guess. (Runway's new Agent mode is quite impressive, doing fairly complex story building from just a short text description of what you want. Not error free obviously, but this was pretty great for a one-shot attempt)
-

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation
By
–
PhyMotion Structured 3D Motion Reward for Physics-Grounded Human Generation
-
AI Model Efficiency: Building a Snake Game with Three-Tier Memory
By
–
Nostalgia hit for the millennials: let's build a Snake Game 🐍
— SambaNova (@SambaNovaAI) 15 mai 2026
Watch what happens when you ask an AI model to build something. Three-tier memory moves models and activations across DDR, HBM, and on-chip SRAM to keep inference fast & efficient. https://t.co/F3zKGeG829 pic.twitter.com/sQfQuJC8G7Nostalgia hit for the millennials: let's build a Snake Game Watch what happens when you ask an AI model to build something. Three-tier memory moves models and activations across DDR, HBM, and on-chip SRAM to keep inference fast & efficient. https://
sambanova.ai/blog/why-dataf
low-matters-more-than-ever?utm_source=x&utm_medium=organic&utm_content=blog-announcement
… -

Technical Survey of Multi-Agent AI Systems and Self-Evolution
By
–
// Beyond Individual Intelligence // One of the more useful multi-agent surveys I've read this year. 200+ papers mapped along three axes: collaboration mechanisms, failure attribution, and self-evolution. The self-evolution chapter is the cleanest field map of where memory,
-

Comparison of Agentic Search vs. Vector Search
By
–
Great paper discussing agentic search vs. vector search.
