Can AI really grasp how articulated objects like robot arms or tools move and change shape? Enter CAPER++: a unified framework for category-level articulated object pose perception. It models objects as root and connected parts linked by joints, uses a clever math trick (SE(3)
@jiqizhixin
-

Spatial-Agent: AI grounded in spatial science for geospatial tasks
By
–
Why do AI agents still struggle with real geospatial tasks like urban planning or disaster response? Researchers from Emory, Rutgers, and UT Austin introduce Spatial-Agent—a new AI that grounds reasoning in spatial science theory instead of web search or pattern matching. It
-

EdgeRazor: AI Models Run on Phones via Mixed-Precision Quantization
By
–
Wow, AI models can run on your phone without losing smarts! Researchers from Nanjing University and Microsoft AI present EdgeRazor — a lightweight framework that uses mixed-precision quantization-aware distillation. It assigns different bit-widths to different parts of the
-

VeRL-Omni: General RL Post-Training Framework
By
–
Cool project! VeRL-Omni is a general RL post-training framework built on verl & vLLM-Omni. Handles heterogeneous pipelines, flexible reward engines, modular backends. Achieves high throughput for image/video/audio gen & understanding. Outperforms existing in efficiency. Code:
-

Flow-OPD distills specialized teachers into one student model
By
–
What if your text-to-image model could master every task without trade-offs? Researchers from USTC, UCLA, CUHK, and Xiaohongshu present Flow-OPD. They distill specialized teachers into one student via on-policy learning, plus a regularizer to protect image quality. Result:
-

Benchmark lets AI agents curate own training data without humans
By
–
What if an AI agent could curate its own training data—without human hand-holding? Researchers from Virginia Tech, UIUC, UW-Madison, and UC Berkeley present CURATION-BENCH, a benchmark that gives generalist coding agents command-line access to inspect, implement, test, and
-

Researchers propose CODA to keep data on chip longer for AI training
By
–
Can AI training be fixed by keeping data on the chip longer? Researchers from MIT, Princeton, Together AI, and Meta introduce CODA — a new way to rewrite Transformer building blocks as GEMM-plus-epilogue programs. Instead of moving large intermediate tensors back and forth to
-

AgentChord: pre-wired recovery for proactive robot task graphs
By
–
What if robots could predict and prevent failures before they happen—instead of just reacting to them? Researchers from CUHK Shenzhen, DexForce, and Shenzhen Loop Area Institute introduce AgentChord: a system that builds a task graph ahead of time, pre-wired with recovery
-

Visual Para-Thinker Uses Parallel Reasoning to Overcome AI Visual Plateaus
By
–
Why do AI models hit a wall in visual reasoning? Researchers from Zhejiang University, Hunan University, and Xiaomi introduce Visual Para-Thinker. It uses parallel divide-and-conquer reasoning to avoid sequential thinking plateaus. Outperforms on V*, CountBench, RefCOCO, and
-

Hallo-Live: Real-Time Joint Audio-Video Avatar Generation
By
–
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation Paper: https://
arxiv.org/abs/2604.23632
Code: https://
github.com/fudan-generati
ve-vision/Hallo-Live
… Our report: https://
mp.weixin.qq.com/s/LCgg_MzjSHqv
YxPIOIhZIw
… #PapersAccepted by Jiqizhixin
