SmolVLM: Redefining small and efficient multimodal models SmolVLM is a family of highly efficient, small-scale vision-language models (VLMs) engineered for low-memory, real-time multimodal inference on mobile and edge devices. These models achieve competitive or even superior
@askalphaxiv
-

VAPO: Value-Based RL Framework for Advanced LLM Reasoning
By
–
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks VAPO is a new value-based reinforcement learning framework designed to enhance long chain-of-thought reasoning in large language models. Built upon Qwen2.5-32B, it outperforms existing methods in
-

SWiRL: Synthetic Data and RL for Language Model Reasoning
By
–
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use This paper introduces Step-Wise Reinforcement Learning (SWiRL), a method for improving multi-step reasoning and tool use in language models through synthetic data generation and offline RL. Problem: Standard RL
-

Test-Time Training Enables One-Minute Video Generation
By
–
One-Minute Generation with Test-Time Training This paper augments Transformers with Test-Time Training (TTT) layers—neural networks used as hidden states—to generate coherent one-minute videos from text storyboards. Problem: Long video generation is bottlenecked by
-

Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
By
–
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention This paper introduces Hogwild! Inference, a parallel LLM inference framework where multiple model instances collaborate by sharing a synchronized attention cache and dynamically adapting their strategies in
-

Top 10 Papers: Video Generation and Multimodal AI Models
By
–
Huge week for video generation and multimodal models, with detailed one-minute video generation and more efficient approaches to multimodality Check out the top 10 papers for the week – Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
– One-Minute -

Test-Time Training enables one-minute coherent video generation
By
–
One-minute video generation—no stitching required Test-Time Training layers can generate long, coherent videos from text storyboards Coherent multi-scene stories with dynamic motion Beats Mamba 2 & DeltaNet by 34 Elo Open code & videos Trending #1 on alphaXiv
-
AlphaXIV Hiring Founding Engineers for Research Platform
By
–
Join Our Team We're hiring founding engineers to help connect the world of researchers. If you have an excellent webdev or mobile background and align with our mission, please DM us or email jobs@alphaxiv.org!
-
New AI Assistant Tool Launch Announcement
By
–
Check out our Assistant here! https://
alphaxiv.org/assistant -
Deep Research for arXiv: AI-Powered Literature Review Tool
By
–
Introducing Deep Research for arXiv
— alphaXiv (@askalphaxiv) 8 avril 2025
Ask questions like 'What are the latest breakthroughs in RL fine-tuning?' and get comprehensive lit reviews with trending papers automatically included
Turn hours of literature searches into seconds with AI-powered research context ⚡ pic.twitter.com/R08xzqbuGyIntroducing Deep Research for arXiv Ask questions like 'What are the latest breakthroughs in RL fine-tuning?' and get comprehensive lit reviews with trending papers automatically included Turn hours of literature searches into seconds with AI-powered research context
