Contextures: The Mechanism of Representation Learning This work introduces contexture theory, a unified mathematical framework that explains what foundation models actually learn during pretraining and why these learned representations generalize to diverse downstream tasks.
@askalphaxiv
-

I-Con: Unified Framework Connects 23 Representation Learning Methods
By
–
I-Con: A Unifying Framework for Representation Learning I-Con introduces a single, unified information-theoretic framework that reveals a deep connection between over 23 different representation learning methods, showing they all minimize a form of KL divergence between
-

WebThinker: Deep Research Agent for Large Reasoning Models
By
–
WebThinker: Empowering Large Reasoning Models with Deep Research Capability WebThinker is a deep research agent that enhances large reasoning models (LRMs) by enabling them to autonomously search the web, navigate pages, and write research reports in real time—overcoming the
-

Prisma: Open Source Mechanistic Interpretability Toolkit for Vision
By
–
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Prisma is an open-source toolkit that brings powerful mechanistic interpretability tools—previously focused on language models—into the vision and video domain, enabling large-scale analysis and
-

TRAJAN: Motion-Centric Evaluation for Generative Video Models
By
–
Direct Motion Models for Assessing Generated Videos This paper introduces TRAJAN, a new motion-centric evaluation method for generative video models that captures temporal inconsistencies more effectively than metrics like Fréchet Distance (FVD), enabling both
-

Mem0: Scalable Long-Term Memory for Production AI Agents
By
–
Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory The paper introduces Mem0, a scalable memory architecture for AI agents that enables long-term conversational coherence by dynamically extracting, consolidating, and retrieving salient information across
-

Chatbot Arena Leaderboard Biases Distort LLM Quality Perception
By
–
The Leaderboard Illusion This study investigates structural biases in Chatbot Arena, a widely used leaderboard for evaluating large language models (LLMs), revealing how selective testing practices and data access asymmetries distort perceptions of model quality. Problem: While
-

X-Fusion Adds Vision to Frozen Language Models
By
–
X-Fusion: Introducing New Modality to Frozen Large Language Models X-Fusion is a dual-tower framework that introduces vision capabilities to pretrained large language models (LLMs) without altering their language weights, enabling unified multimodal understanding and generation.
-

Reinforcement Learning Enhances LLM Mathematical Reasoning
By
–
Reinforcement Learning for Reasoning in Large Language Models with One Training Example This paper demonstrates that Reinforcement Learning with Verifiable Reward (RLVR) using just one training example (1-shot RLVR) can significantly enhance the mathematical reasoning abilities
-

LLM Leaderboards Gaming: Chatbot Arena Scoring Integrity Questioned
By
–
Are LLM leaderboards no longer trustworthy? @cohere
's deep dive reveals how Chatbot Arena scores can be gamed Providers test 10–27 private models & submit the best Proprietary models receive 2–3× more data Rankings may reflect overfitting Trending on alphaXiv
