DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments DeepResearcher is a reinforcement learning framework that trains LLM agents to perform deep research through real-world web interactions, moving beyond static prompt or RAG-based
LLMS
-

Language Model Size and Reasoning Capability Scaling Laws
By
–
Do Larger Language Models Imply Better Reasoning? A Pretraining Scaling Law for Reasoning This paper investigates the relationship between the size of language models (LLMs) and their reasoning abilities, focusing on a synthetic multi-hop reasoning task based on real-world
-

VAPO: Value-Based RL Framework for Advanced LLM Reasoning
By
–
VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks VAPO is a new value-based reinforcement learning framework designed to enhance long chain-of-thought reasoning in large language models. Built upon Qwen2.5-32B, it outperforms existing methods in
-

SWiRL: Synthetic Data and RL for Language Model Reasoning
By
–
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use This paper introduces Step-Wise Reinforcement Learning (SWiRL), a method for improving multi-step reasoning and tool use in language models through synthetic data generation and offline RL. Problem: Standard RL
-

Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
By
–
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention This paper introduces Hogwild! Inference, a parallel LLM inference framework where multiple model instances collaborate by sharing a synchronized attention cache and dynamically adapting their strategies in
-

Top 10 Papers: Video Generation and Multimodal AI Models
By
–
Huge week for video generation and multimodal models, with detailed one-minute video generation and more efficient approaches to multimodality Check out the top 10 papers for the week – Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
– One-Minute -
SE2 and SE1 AI Model Updates: Smaller Releases Incoming
By
–
Those SE2 updates (VS 1.1 and upcoming VS 1.2) are much smaller than the usual SE1 updates. Also, the next SE1 update is very near.
-
Google offers free prompt engineering learning here
By
–
Learn prompt engineering here for free by Google. Impressive.
-
Google Ironwood TPU 7th Gen Optimized for LLM Training
By
–
L’Ironwood TPU est une claque technologique.
— VISION IA (@vision_ia) 12 avril 2025
C’est la 7e génération de puces IA de Google, conçue sur-mesure pour l’ère des LLM. Plus rapide, plus efficace, plus optimisée que les GPU traditionnels.
On parle d’une infrastructure taillée pour entraîner et faire tourner des IA… pic.twitter.com/rq4daqq8RSL’Ironwood TPU est une claque technologique. C’est la 7e génération de puces IA de Google, conçue sur-mesure pour l’ère des LLM. Plus rapide, plus efficace, plus optimisée que les GPU traditionnels. On parle d’une infrastructure taillée pour entraîner et faire tourner des IA
-

Sample-Based Approach Improves Language Model Test-Time Alignment
By
–
Sample, Don't Search: Rethinking Test-Time Alignment for Language Models Gonçalo Faria, Noah A. Smith: https://
arxiv.org/abs/2504.03790 #ArtificialIntelligence #DeepLearning #MachineLearning
