What if LLM reinforcement learning could assign credit more accurately by thinking step by step? Researchers from Peking University and Microsoft Research Asia introduce GenAC: a generative critic that replaces one-shot value predictions with chain-of-thought reasoning before
RESEARCH
-

GPT-5.5 Pro applies research technique to generate funny word pairs
By
–

GPT-5.5 Pro faces its hardest academic challenge: to apply the technique from a paper analyzing which word pairs were funny & why to come up with its own It came up with scrotum snorkel, tuba subpoena, waffle coffin, toad commode, diarrhea tiara, banana tribunal & muffin ruffian
-

Top AI Papers of the Week (May 11–17)
By
–
The Top AI Papers of the Week (May 11 – May 17) – AEvo
– δ-mem
– AutoTTS
– AI Co-Mathematician
– Lighthouse Attention
– Is Grep All You Need?
– A Geometric Calculator Inside a Neural Network Read on for more: -

StarVLA: Modular Vision-Language Robot Codebase
By
–
What if building a robot that sees, understands, and acts was as easy as snapping Lego together? Enter StarVLA: a modular codebase that lets you swap vision-language or world-model backbones and action heads independently. It matches or surpasses prior methods on benchmarks
-

Causal vs Observational Agency in Agents
By
–

Causal vs observational agency. Agent actions (a) should be treated as interventions, not as evidence for hypotheses (p). Actions by other agents or tool outputs are evidence (o).
-

One-line fix to prevent LLM agent delusions
By
–
One line of code is all it takes to prevent LLM agent delusions, instead of post-training patches like RL. https://
love4all.ai/blog/why-it-is
-important-to-understand-causality-and-agency/
… 4 ∀ https://
github.com/nandodef/love4
all-ai/tree/main/docs/files
… -

SpaceXAI: Grok next version trained on 1.5T V9 model, upgrade coming summer
By
–

SPACEXAI : The next version of Grok, based on the 1.5T V9 base model has finished training. Looks like we will get a major upgrade this summer. > Next, we are adding the Cursor data in supplemental training. Soon
-
Improvement of Grok with V9 training and Cursor data
By
–
We are improving the base Grok 0.5T V8 model (public version 4.3) every few days. The 1.5T V9 has just completed its training (incorrectly called pre-training) and represents a major upgrade. Then, we add Cursor data into a
-
Study finds memory in LLM agents remains unreliable
By
–
Breaking new study: memory in LLM agents still can’t really be trusted, even after over trillion dollars has gone into the development of the field.
-
Reinforcement-Learned Atlas Performs Human Mocap-Inspired Moves
By
–
Reinforcement-Learned Atlas Performs Moves Inspired by Human Mocap and Animation
— Ronald van Loon (@Ronald_vanLoon) 17 mai 2026
via @ZappyZappy7
#Robotics #Engineering #ArtificialIntelligence #Innovation #Technology pic.twitter.com/Udic9QKJTIReinforcement-Learned Atlas Performs Moves Inspired by Human Mocap and Animation
via @ZappyZappy7 #Robotics #Engineering #ArtificialIntelligence #Innovation #Technology
