Cambrian-S: Towards Spatial Supersensing in This paper boasts an impressive roster of advisors, including Rob Fergus, Yann LeCun, Fei-Fei Li, and Saining Xie, and they aim to answer this question: can AI go beyond “seeing” to truly understanding space? They propose
@jiqizhixin
-

Scaling Agent Learning Through Experience Synthesis
By
–
Scaling Agent Learning via Experience Synthesis Paper: https://
arxiv.org/abs/2511.03773 -

DreamGym: Training LLM Agents Through Synthetic Reasoning
By
–
Can LLM agents learn by dreaming? DreamGym from Meta is a new framework that lets AI agents train via synthetic reasoning-based experiences instead of costly real rollouts. It models environment dynamics, replays and adapts tasks, and even improves sim-to-real transfer.
-

ReasonMed: 370k Medical Reasoning Dataset Advances Clinical AI
By
–
ReasonMed: the largest medical reasoning dataset, advancing LLM performance in clinical QA. Comprising 370k curated examples distilled from 1.75M reasoning paths, ReasonMed is built through a multi-agent EMD (easy–medium–difficult) pipeline with generation, verification, and an
-

High School Geometry Improves Spatial Intelligence in AI Models
By
–
Can high school geometry teach AI to understand space? A new study tackles the critical challenge of spatial intelligence in Multimodal Large Language Models (MLLMs). Researchers found that fine-tuning models on Euclid30K, a new dataset of ~30,000 Euclidean geometry
-

RemeDi: LLMs Self-Correct Mistakes Through Remasking Mechanism
By
–
LLMs could spot their own mistakes and revise them on the fly. Researchers are tackling this with RemeDi (Remasking-enabled Diffusion Language Model), a new model that introduces a "remasking" mechanism for more flexible text generation. Instead of just predicting tokens,
-

AI Search Agents Vulnerable to Malicious Website Manipulation
By
–
How safe are AI search agents when browsing the real, messy internet? A new study reveals that LLM-powered search agents are highly vulnerable to being misguided by unreliable or malicious websites. To test this, researchers developed SafeSearch, an automated red-teaming
-
RoboChallenge: Embodied AI Reaches ImageNet Moment
By
–
Embodied AI has reached its ImageNet moment.
— 机器之心 JIQIZHIXIN (@jiqizhixin) 19 octobre 2025
Dexmal, in collaboration with Hugging Face, has released RoboChallenge.
It's the first large-scale, multi-task benchmark in which real robots perform manipulation tasks in real-world physical environments.
Our report:… pic.twitter.com/UVBZdUXUjhEmbodied AI has reached its ImageNet moment. Dexmal, in collaboration with Hugging Face, has released RoboChallenge. It's the first large-scale, multi-task benchmark in which real robots perform manipulation tasks in real-world physical environments. Our report:
-
EgoAgent: First-Person AI Perception and Action Model
By
–
How can we build AI agents that perceive, predict, and act from a first-person view, just like humans?
— 机器之心 JIQIZHIXIN (@jiqizhixin) 18 octobre 2025
Inspired by the human perception-action loop, researchers propose EgoAgent, a unified agent model that simultaneously learns to represent the environment, predict the future,… pic.twitter.com/Gd4JWaZBTRHow can we build AI agents that perceive, predict, and act from a first-person view, just like humans? Inspired by the human perception-action loop, researchers propose EgoAgent, a unified agent model that simultaneously learns to represent the environment, predict the future,
-

PaDT: MLLMs Generate Visual Detection Outputs Directly
By
–
Ever wonder if an AI could do more than just describe an image and actually show you where things are? PaDT (Patch-as-Decodable Token) is a unified paradigm enabling Multimodal Large Language Models (MLLMs) to directly generate visual outputs like detection boxes and
