Sometimes, a single LLM isn’t enough, and we could coordinate multiple models to solve complex tasks together. Router-R1 is a reinforcement learning–based framework that routes and aggregates multiple LLMs like an intelligent conductor. Key ideas: – Formulates multi-LLM
@jiqizhixin
-

Entropy-Balanced Policy Optimization for AI Agents
By
–
Agentic Entropy-Balanced Policy Optimization Renmin University of China, Kuaishou Technology
Paper: https://
huggingface.co/papers/2510.14
545
…
Code: https://
github.com/dongguanting/A
RPO
… -

VIR-Bench: Evaluating Multimodal LLMs on Travel Video Understanding
By
–
How well can multimodal LLMs understand long-distance travel videos? Enter VIR-Bench, a new benchmark with 200 real-world travel videos that challenges models to reconstruct itineraries and reason over extended geospatial-temporal trajectories. Why it matters: mastering
-

RiskPO: Risk-Based Policy Optimization for LLM Post-Training
By
–
#PapersAccepted by Jiqizhixin
Our report: https://
mp.weixin.qq.com/s/9TbUIT6ed_wO
viVU0GuLqg
… RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training Peking University
Paper: https://
arxiv.org/abs/2510.00911
v1
…
Code: https://
github.com/RTkenny/RiskPO -

Risk-based Policy Optimization Improves LLM Reasoning Through Reward Risk-Taking
By
–
RL keeps evolving! Now you can teach LLMs to reason better by rewarding risk-taking. Risk-based Policy Optimization (RiskPO) is a new reinforcement learning framework for post-training LLMs. Instead of averaging rewards like GRPO, RiskPO uses a Mixed Value-at-Risk objective
-

Conditional Representation Learning for Customized AI Tasks
By
–
#PapersAccepted by Jiqizhixin
Our report: https://
mp.weixin.qq.com/s/krvT5dViY2x0
TbPKomGThg
… Conditional Representation Learning for Customized Tasks Sichuan University, Chinese Academy of Sciences
Paper: https://
arxiv.org/abs/2510.04564
Code: https://
github.com/XLearning-SCU/
2025-NeurIPS-CRL
… -

Conditional Representation Learning: Tailoring AI Embeddings to User Preferences
By
–
How can we make AI representations adapt to what humans actually care about? This paper introduces Conditional Representation Learning (CRL), a new paradigm that tailors image embeddings to user-specified criteria. Key idea: – The semantics of a feature space are defined by
-

RAPID Hand: Robust Dexterous Robot Manipulation Platform
By
–
#PapersAccepted by Jiqizhixin
Our report: https://
mp.weixin.qq.com/s/x-pov_ppBKXw
Qv_3NmMquw
… RAPID Hand: A Robust, Affordable, Perception-Integrated, Dexterous Manipulation Platform for Generalist Robot Autonomy Sun Yat-sen University, University of California, Merced, CASIA
Paper: -
RAPID Hand: Dexterous Affordable Generalist Robot Platform
By
–
How can we make generalist robot hands both dexterous and affordable?
— 机器之心 JIQIZHIXIN (@jiqizhixin) 17 octobre 2025
RAPID Hand is a co-designed hardware & software platform with:
– 20-DoF compact robotic hand
– Wrist vision + fingertip tactile + proprioception (sub-7 ms latency)
– High-DoF teleoperation with stable… pic.twitter.com/ETui9LHQCGHow can we make generalist robot hands both dexterous and affordable? RAPID Hand is a co-designed hardware & software platform with: – 20-DoF compact robotic hand
– Wrist vision + fingertip tactile + proprioception (sub-7 ms latency)
– High-DoF teleoperation with stable -
Gaussian Embeddings: How JEPAs Learn Data Density
By
–
Gaussian Embeddings: How JEPAs Secretly Learn Your Data Density Paper:
