Trending AI papers on alphaXiv this week, featuring an incredible week for the computer vision and multimodal communities – Navigation World Models (discussion with author @_amirbar
)
– Open-Sora Plan: Open-Source Large Generation Model (discussion with author
@askalphaxiv
-

Latest AI Papers: Navigation Models and Video Generation
By
–
-

Multi-Hop Reasoning Limitations in Large Language Models
By
–
Evaluating Latent Multi-Hop Reasoning in Large Language Models Investigates whether LLMs can combine factual knowledge across multiple hops without relying on pretraining shortcuts. Problem:
LLMs struggle with latent multi-hop reasoning due to reliance on pretraining -

Make-It-Animatable: Automated 3D Character Rigging Framework
By
–
Make-It-Animatable Efficient framework for converting any 3D humanoid character into animation-ready form in under one second. Problem:
Rigging and skinning 3D characters are labor-intensive and existing methods fail to generalize across diverse shapes and poses. Method: -

Enhancing LLM Reasoning with Separate Critique Model
By
–
Enhancing LLM Reasoning Improving LLM reasoning by using a separate "critique" model to provide feedback during both training and testing. Problem:
LLMs struggle with complex reasoning and self-correction, especially on challenging tasks where performance plateaus. Method: -

LLM-Powered GUI Agents: Automation Survey and Insights
By
–
Large Language Model-Brained GUI Agents: A Survey A comprehensive survey exploring how large language models can be used to create AI agents that interact with graphical user interfaces through natural language. Problem:
GUI automation lacks flexibility, requiring complex -

Unified Diffusion Model for Image Generation and Understanding
By
–
One Diffusion to Generate Them All A unified diffusion model that can perform both image generation and understanding tasks through a simple sequence-based approach. Problem:
Diffusion models are task-specific, lacking the universality and flexibility seen in large language -

Model Distillation and O1 Replication: Research Ethics Concerns
By
–
O1 replication journey (part 2) A critical examination of how simple model distillation can match O1's mathematical abilities, raising concerns about AI research practices. Problem: Replicating O1 model capabilities has encouraged the use of undisclosed distillation
-

ShowUI: Lightweight Vision-Language Model for GUI Automation
By
–
ShowUI A lightweight vision-language model for efficient GUI automation that achieves SOTA performance while using fewer parameters and less training data. Problem:
GUI automation relying on APIs and metadata struggles with real-world visual scenarios requiring efficient, -

Critic-V: External Feedback Framework for Vision-Language Models
By
–
Critic-V A framework using an external critic model to refine vision-language model outputs through feedback. Problem: Vision-language models struggle with errors like hallucinations and irrelevant reasoning due to lack of external feedback. Method: Combines a Reasoner that
-

Janus Framework Unifies Multimodal AI Understanding and Generation
By
–
Introducing Janus, a novel framework that
unifies multimodal understanding and generation. By decoupling visual encodings, Janus can easily extend to new input types such as audio and EEG signals Excited to have the author @wu_chengyue from @deepseek_ai to discuss their work!