Multimodal LLMs often "overthink"—producing long, verbose answers even for simple visual questions FAST (Fast-Slow Thinking for LVLMs) achieves SOTA accuracy while slashing token usage +10% accuracy over baselines Up to 67% fewer tokens used Trending on alphaXiv
@askalphaxiv
-

TTRL Boosts LLM Reasoning by 159% Using Unlabeled Test Data
By
–
LLMs usually rely on labeled data to improve But what if they could self-improve at test time — without any labels? Introducing TTRL: a new RL framework that boosts LLM reasoning by up to 159% on the AIME 2024 using just unlabeled test data Trending #1 on alphaXiv
-

LUFFY Framework Boosts LLM Math Reasoning by Seven Points
By
–
LLMs still struggle with deep reasoning LUFFY is a new framework bridging imitation and exploration by injecting off-policy guidance (like DeepSeek-R1) into zero-RL Boosts math reasoning by 7 points Improves weaker base models like Llama and Qwen Trending on alphaXiv
-

ReDi Framework Improves Image Generation via Joint Feature Synthesis
By
–
Boosting Generative Image Modeling via Joint Image-Feature Synthesis This paper introduces ReDi, a new framework that jointly models low-level VAE latents and high-level semantic features within the diffusion process, significantly improving image generation quality and training
-

LUFFY: Off-Policy Reasoning for Zero-Shot Reinforcement Learning
By
–
Learning to Reason under Off-Policy Guidance This paper presents LUFFY (Learning to reason Under oFF-policY guidance), a framework that improves zero-shot reinforcement learning (zero-RL) by incorporating off-policy reasoning traces alongside on-policy exploration, enabling
-

Data Scaling Laws for End-to-End Autonomous Driving Systems
By
–
Data Scaling Laws for End-to-End Autonomous Driving This paper explores how scaling training data affects performance in end-to-end autonomous driving systems, evaluating data efficiency and scaling laws with a unified, differentiable model. Problem: Modular autonomous driving
-

VideoChat-R1: Spatio-Temporal Perception via Reinforcement Fine-Tuning
By
–
Chat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning This paper explores the use of Reinforcement Fine-Tuning (RFT) with Group Relative Policy Optimization (GRPO) to enhance spatio-temporal perception in video multimodal large language models (MLLMs),
-

DeepResearcher: Scaling Deep Research via Reinforcement Learning
By
–
DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments DeepResearcher is a reinforcement learning framework that trains LLM agents to perform deep research through real-world web interactions, moving beyond static prompt or RAG-based
-

Language Model Size and Reasoning Capability Scaling Laws
By
–
Do Larger Language Models Imply Better Reasoning? A Pretraining Scaling Law for Reasoning This paper investigates the relationship between the size of language models (LLMs) and their reasoning abilities, focusing on a synthetic multi-hop reasoning task based on real-world
-

MedSAM2: Foundation Model for 3D Medical Image Segmentation
By
–
MedSAM2: Segment Anything in 3D Medical Images and Videos MedSAM2 is a foundation model for 3D medical image and video segmentation, extending SAM2 to support volumetric and temporal data across diverse clinical tasks. Problem: Existing segmentation models are mostly limited to
