Token-Budget-Aware LLM Reasoning The paper proposes a framework for improving the efficiency of reasoning in large language models (LLMs) by dynamically estimating and adjusting token budgets based on problem complexity. This approach helps reduce the token overhead caused by
@askalphaxiv
-

Thousand Brains Project: Neocortex-Based AI Paradigm
By
–
The Thousand Brains Project: A New Paradigm for Sensorimotor Intelligence The Thousand Brains Project aims to create AI systems based on the principles of the neocortex, focusing on sensorimotor learning and interaction with the environment to build world models, similar to
-

FaceLift: Single Image 3D Head Reconstruction with Diffusion
By
–
FaceLift: Single Image to 3D Head with View Generation and GS-LRM FaceLift is a two-stage pipeline for high-quality 3D head reconstruction from a single image. It uses a diffusion model for multi-view generation and a GS-LRM model to create a detailed 3D head representation,
-

DRT-o1: Deep Reasoning Translation Optimizes Neural Machine Translation
By
–
DRT-o1: Optimized Deep Reasoning Translation via Long Chain-of-Thought Problem: Neural machine translation struggles with translating texts containing similes or metaphors, often failing to convey the intended meaning due to cultural differences and the limitations of literal
-

Joint 3D Reconstruction of Humans, Scenes, and Cameras
By
–
Reconstructing People, Places, and Cameras A method that jointly reconstructs human meshes, scene point clouds, and camera parameters from sparse, uncalibrated images, integrating human pose and scene geometry for improved world-scale 3D reconstruction. Problem: Existing scene
-

LearnLM: Gemini Enhanced for Personalized AI Tutoring
By
–
LearnLM: Improving Gemini for Learning An enhanced model, LearnLM, is designed to improve AI tutoring by training Gemini models to follow pedagogical instructions and adapt to diverse educational contexts. Problem: Existing generative AI systems lack tailored pedagogical
-

ResearchTown: Simulating Human Research Communities with Agent-Data Graphs
By
–
ResearchTown: Simulator of Human Research Community The framework models research activities through an agent-data graph, with researchers as agents and papers as data, utilizing a TextGNN mechanism for message-passing between nodes to simulate activities like paper writing and
-

Mulberry: MLLM Reasoning via Collective Monte Carlo Tree Search
By
–
Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search The paper presents Mulberry, an MLLM with step-by-step reasoning and reflection, powered by CoMCTS. CoMCTS uses collective knowledge from multiple models to efficiently
-

Align-Anything: Training Multimodal Models with Language Feedback
By
–
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback The paper introduces the Align-Anything framework to enhance all-modality models' alignment with human preferences using language feedback, improving instruction-following across text,
-

Reproducing o1: Reinforcement Learning Search Scaling Roadmap
By
–
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective This paper presents a roadmap for reproducing OpenAI's o1, an advanced LLM that excels in reasoning. It focuses on four key components—policy initialization, reward design, search,
