Can one diffusion model master multiple text-to-image tasks without forgetting? Researchers from Fudan University & Alibaba Group present DiffusionOPD. Instead of joint training, they first train separate expert teachers, then distill their knowledge into a single student
@jiqizhixin
-

OSCAR: 2-bit KV cache for LLMs without accuracy loss
By
–
Can LLMs run on ultra-low-bit memory without tanking accuracy? Researchers from Together AI, University of Sydney, and UIUC present OSCAR — a method that uses offline, attention-aware covariance analysis to design fixed rotations and clipping thresholds for 2-bit KV cache
-
Legato: blending actions with noise for smooth robot motion
By
–
What if your robot could move without awkward jerks or hesitation?
— 机器之心 JIQIZHIXIN (@jiqizhixin) 9 juin 2026
Researchers from Shanghai Jiao Tong University and Spirit AI introduce Legato.
It trains vision-language-action models to blend known actions with noise during denoising, making chunk boundaries naturally… pic.twitter.com/MP9OMwszhfWhat if your robot could move without awkward jerks or hesitation? Researchers from Shanghai Jiao Tong University and Spirit AI introduce Legato. It trains vision-language-action models to blend known actions with noise during denoising, making chunk boundaries naturally
-

Generative compression shrinks Earth observation data 10,000x without loss
By
–
Can Earth observation data be shrunk 10,000x without losing scientific value? Tsinghua, Sun Yat-Sen, and National University of Singapore researchers present a generative compression model that learns from historical Earth archives. Instead of treating compression as a
-

ThoughtTrace dataset reveals users’ real thoughts during AI chats
By
–
Ever wonder what users are really thinking during AI conversations? Researchers from Johns Hopkins, MIT, and Google Research introduce ThoughtTrace—the first large-scale dataset pairing real-world user-AI chats with people’s own thoughts: why they sent a prompt and how they
-

Orbit: RL-train trillion-parameter LLMs on just 8 GPUs
By
–
What if you could RL-train trillion-parameter LLMs on just 8 GPUs? Enter Orbit, an open-source framework that keeps the base model fixed and trains only a tiny BF16 adapter. It outperforms standard sync RL: 71% faster step times, 50% higher rollout throughput, 81% less train
-

TextPro-SLM Approach Bridges Speech and Text AI Gap
By
–
Why do speech AI models still lag behind text AI? Researchers at CUHK and Huawei propose TextPro-SLM — an approach that shrinks the gap by making spoken input look more like text input. Instead of tweaking the output, they redesign the input side with a unified speech
-

Samsung MeKi: Memory-Based Expert Knowledge Injection for LLMs
By
–
Want to scale LLMs without skyrocketing compute costs? Samsung presents MeKi: Memory-based Expert Knowledge Injection. Instead of making models bigger to learn everything, MeKi gives LLMs a dynamic memory bank of expert knowledge. Think of it as a cheat sheet the model can
-

LA-Pose: Self-driving car learns position from unlabeled driving videos
By
–
What if a self-driving car could learn its position just by watching millions of hours of driving video? Researchers at Wayve and Simon Fraser University introduce LA-Pose: they train a model to learn "latent actions"—hidden motion patterns—from unlabeled driving footage, then
-

PhoneHarness: Seamlessly mixing CLI, GUI, and host tools for phone agents
By
–
Can your phone agent do more than just tap and swipe? Tencent Hunyuan researchers (with CUHK & Tsinghua) introduce PhoneHarness. It’s a new harness that lets phone agents seamlessly mix CLI, GUI, and host tools—verifying real side effects, not just screen predictions. Result:
