"OPRD: On-Policy Representation Distillation" On-policy distillation usually matches teacher and student only at the token probability level, throwing away the teacher’s hidden states. This paper moves the loss before the LM head, aligning student and teacher representations on
@askalphaxiv
-

Self-Trained Verification for Self-Improvement in Training and Test Phases
By
–
Reasoning models improve faster with a good verifier, but verifiers cannot learn to detect subtle errors by themselves. However,
-

Critique of AI benchmarks not measuring real work
By
–
"The Final Exam of Agents" With the way AI continues to brilliantly succeed at benchmarks, yet this has not translated into real economic value, the authors of this article say the problem lies with the benchmarks, since none measure work.
-

Microsoft shares deep dive into data engineering for frontier AI models
By
–
"MAI-Thinking-1: Building a Hill-Climbing Machine" Microsoft just did something almost no frontier AI lab has done before They shared how they engineered the data behind a frontier-scale model in unusual depth. From data collection and eval decontamination, to data mix
-

Trust Region On-Policy Distillation: Learning from reliable teacher signals
By
–
“Trust Region On-Policy Distillation” On-policy distillation is powerful, but one bad mismatch between student and teacher can negatively impact the gradients. So this paper's TrOPD only learns where the teacher is reliable, treats outliers separately, and nudges the student
-

New Benchmark for Visual State Tracking in Video Understanding
By
–
"Benchmarking Visual State Tracking in Multimodal Understanding" A new benchmark for tracking visual states. Even though video MLLMs can describe clips really well, they still cannot reliably track what changes over time. This benchmark contains 834 videos and 1,500
-

Scaling PEFT: Towards Million Personal Models of Trillion Parameters
By
–
"On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters" Right now LLM personalization mostly means prompts, memory, or retrieval on top of one shared assistant. This paper instead keeps one trillion-parameter base model shared, and give each user a tiny
-

VLMs learn 3D natively, skipping expert architectures and complex designs
By
–
"VLM^3: VLMs Are Native 3D Learners" This paper shows that VLMs can learn 3D natively. Most 3D vision systems rely on expert architectures, regression heads, heavy augmentations, and task-specific losses. But they show that you can skip the majority of these designs. All they
-
AlphaXiv.org now offers GPT 5.5 and Claude models
By
–
Now available in addition to GPT 5.5 and Claude. Check out http://
alphaXiv.org! -
Gemini 3.5 Flash: Understand Research Papers Faster with Highlighted Q&A and Cross-Paper Context
By
–
Introducing Gemini 3.5 Flash for understanding research papers 🚀
— alphaXiv (@askalphaxiv) 2 juin 2026
Highlight any section of a paper to ask questions and “@” other papers for quick context, comparisons, and benchmark references pic.twitter.com/6bmBR63SegIntroducing Gemini 3.5 Flash for understanding research papers Highlight any section of a paper to ask questions and “@” other papers for quick context, comparisons, and benchmark references
