1. Phi-4-Mini-Reasoning Microsoft released Phi-4-Mini-Reasoning to explore small reasoning language models for math.
@dair_ai
-
Top AI Papers of the Week: April 28 – May 4
By
–
Here are the top AI Papers of the Week (April 28 – May 4): – Kimi-Audio
– UniversalRAG
– LLM for Engineering
– DeepSeek-Prover-V2
– Phi-4-Mini-Reasoning
– Advances and Challenges in Foundation Agents Read on for more: -

General-Reasoner: Reinforcement Learning Boosts LLM Reasoning
By
–
9. General-Reasoner General-Reasoner is a reinforcement learning approach that boosts LLM reasoning across diverse domains by using a 230K-question dataset and a model-based verifier trained to understand semantics beyond exact matches.
-
Tiny Reasoning Models: 1.5B Parameter Tina Family
By
–
10. Tiny Reasoning Models Tina is a family of 1.5B parameter reasoning models trained using LoRA-based reinforcement learning (RL) to achieve high reasoning accuracy at very low cost.
-

Assessing LLM Goal-Directedness: New Evaluation Framework
By
–
8. Evaluate the Goal-Directedness of LLMs Introduces a new framework to assess whether LLMs use their capabilities effectively toward achieving given goals.
-
NoProp: Gradient-Free Learning Method for Neural Networks
By
–
10). NoProp NoProp is a novel gradient-free learning method where each neural network layer independently learns to denoise a noisy version of the target, inspired by diffusion and flow matching.
-
One-Minute Video Generation with Test-Time Training Innovation
By
–
9). One-Minute Video Generation with Test-Time Training
— DAIR.AI (@dair_ai) 13 avril 2025
One-Minute Video Generation with Test-Time Training introduces TTT layers, a novel sequence modeling component where hidden states are neural networks updated via self-supervised loss at test time.https://t.co/gC78dOaotT9). One-Minute Generation with Test-Time Training One-Minute Generation with Test-Time Training introduces TTT layers, a novel sequence modeling component where hidden states are neural networks updated via self-supervised loss at test time.
-
KnowSelf Framework Enables Dynamic Self-Aware LLM Agents
By
–
8). Agentic Knowledgeable Self-awareness KnowSelf is a new framework that introduces agentic knowledgeable self-awareness, enabling LLM agents to dynamically decide when to reflect or seek knowledge based on situational complexity, mimicking human cognition.
-
Computer Agent Arena: Benchmarking LLM and VLM Agents
By
–
7). Compute Agent Arena
— DAIR.AI (@dair_ai) 13 avril 2025
Computer Agent Arena is a new open platform for benchmarking LLM and VLM-based agents on real-world computer-use tasks, like coding, editing, and web navigation, using a virtual desktop environment.https://t.co/hpaFLGlg2i7). Compute Agent Arena Computer Agent Arena is a new open platform for benchmarking LLM and VLM-based agents on real-world computer-use tasks, like coding, editing, and web navigation, using a virtual desktop environment.
-

LightPROF: Efficient Knowledge Graph Reasoning for Small LLMs
By
–
6). Efficient KG Reasoning for Small LLMs LightPROF is a lightweight framework that enables small-scale language models to perform complex reasoning over knowledge graphs (KGs) using structured prompts.
