Now available in addition to Gemini, Claude, and o3-mini. Check out http://
alphaXiv.org! Alternatively replace the “arXiv” with “alphaXiv” in any arXiv URL. Ex: https://
arxiv.org/abs/2501.17161 -> https://
alphaXiv.org/abs/2501.17161
@askalphaxiv
-
alphaXiv launches AI-powered arXiv paper access tool
By
–
-

NVIDIA Open-Sources GR00T N1 Foundation Model for Humanoid Robots
By
–
Meet GR00T N1: NVIDIA's foundation model for humanoid robots—now fully open-sourced Vision-Language-Action model Trained with real-robot trajectories, human videos, and synthetic data Outperforms SOTA imitation learning baselines Trending on alphaXiv
-

AnyCalib: Model-Agnostic Single-View Camera Calibration Method
By
–
AnyCalib: On-Manifold Learning for Model-Agnostic Single-View Camera Calibration AnyCalib is a model-agnostic method for calibrating camera intrinsic parameters from a single in-the-wild image. It bypasses the need for extrinsic cues and works across multiple camera models,
-

RWKV-7 Goose: State-of-the-art Multilingual Language Model
By
–
RWKV-7 "Goose" with Expressive Dynamic State Evolution RWKV-7 "Goose" is a sequence modeling architecture that achieves state-of-the-art multilingual performance with 3 billion parameters. Despite training on fewer tokens than other top models, it excels in both English and
-

New Metric Measures AI Long-Task Completion Ability
By
–
Measuring AI Ability to Complete Long Tasks This paper introduces a new metric, the 50%-task-completion time horizon, to measure AI's ability to complete long tasks. It tracks the time it takes for AI models to match the 50% success rate in tasks that humans typically complete.
-

SynCity: Training-Free 3D World Generation from Text
By
–
SynCity: Training-Free Generation of 3D Worlds This paper presents SynCity, a novel approach to generating large-scale 3D worlds from textual descriptions. It combines the geometric precision of pre-trained 3D generative models with the artistic flexibility of 2D image
-

Aligning Multimodal LLMs with Human Preferences: Survey
By
–
Aligning Multimodal LLM with Human Preference: A Survey This paper reviews alignment algorithms for Multimodal Large Language Models (MLLMs), which handle tasks involving text, visual, and auditory data, aiming to address issues like truthfulness, safety, and alignment with
-

4DGS-1K Achieves 1000+ FPS Dynamic Scene Rendering Innovation
By
–
1000+ FPS 4D Gaussian Splatting for Dynamic Scene Rendering This paper presents 4DGS-1K, a faster and more memory-efficient framework for dynamic scene rendering, achieving over 1000 FPS on modern GPUs by improving 4D Gaussian Splatting (4DGS). Problem: 4DGS faces two
-

Causal Emergence 2.0: Quantifying Emergent Complexity
By
–
Causal Emergence 2.0: Quantifying emergent complexity This paper presents Causal Emergence 2.0 (CE 2.0), a theory that quantifies how different scales in complex systems contribute to causation and introduces a new measure of emergent complexity based on distributed causal
-
Top 10 Foundation Model Papers: Humanoid Robots and Multimodal LLMs
By
–
This week was huge for foundation models, from generalist humanoid robots to multimodal LLMs that learn from human preferences and negative examples – here are the top 10 papers for the week – DAPO: An Open-Source LLM Reinforcement Learning System at Scale
– GR00T N1: An