« Fixed Point Reasoners: Stable and Adaptive Deep Looped Transformers » While Looped Transformers can devote more depth to harder problems, they still need a good method to know when to stop. This article makes of
@askalphaxiv
-

Do as I Do: Dexerous Manipulation Data from Everyday Human Videos
By
–
“Do as I Do: Dexterous Manipulation Data from Everyday Human Videos” With how robot dexterity is bottlenecked by data as teleoperation and MoCap are expensive and internet videos are only observational, this paper turns normal RGB human videos into executable robot hand
-

Research on anthropomorphic misalignment requires stronger evidence
By
–
« Position: Research on anthropomorphic misalignment needs stronger evidence » AI safety papers often claim that models are deceptive, calculating, or concerned with self-preservation, but many results do not show
-

Q-Learning by Inversion for Flow Policies in Offline RL
By
–
Q-Learning by Inversion. Flow policies should be excellent for offline RL, but their iterative generation of actions makes RL training painful. This paper transforms each flow step into an RL step, then uses inversion to recover
-

Making Depth Reusable by Looping World Models
By
–
“World Models in a Loop” World models require deep computations for stable long-horizon simulation, but deeper models are costly and errors accumulate over deployments. This paper makes depth reusable by looping
-
Research agents excel at resolving implementation issues, not novel discoveries
By
–
1/N Model capabilities today are subpar for end-to-end autonomous research, what we would categorize as the ability to make “novel discoveries”. However, in our experimentation, research agents have proven to be an excellent way to resolve implementation issues and carry out
-
Presentation of auto-search for arXiv articles via autoarxiv
By
–
Introducing autoresearch for arXiv papers
— alphaXiv (@askalphaxiv) 18 juin 2026
Change 'arxiv' to 'autoarxiv' in any paper URL
An agent deploys to resolve setup issues on the codebase, run a minimal reproduction, and estimate full replication cost. Read more below pic.twitter.com/2UHVqbT7EuPresentation of auto-search for arXiv articles. Replace 'arxiv' with 'autoarxiv' in any article URL. An agent deploys to resolve configuration problems in the codebase, run a minimal reproduction, and estimate the replication cost.
-

Variable-Width Transformers: Wide at Start and End, Narrow in the Middle
By
–
This article shows that width should be allocated unevenly, with models wide at the beginning and end but narrow in the middle. Thus, the bottleneck forces better use of representations instead of wasting the
-

Distribution of thought paths in a continuous latent space
By
–
Latent Thought Flow This article moves reasoning into a continuous latent space, but instead of learning a single hidden thought path, it learns a distribution over many paths. Using a continuous GFlowNet, Latent Thought Flow assigns a
-

ExpRL uses reference solutions as reward scaffolds for exploratory RL
By
–
“ExpRL: Exploratory RL for LLM Mid-Training” Sparse reward RL works only when the base model can already find useful reasoning paths, but on hard problems it often gets no signal. This paper uses reference solutions as reward scaffolds instead of imitation targets, letting an
