DR-Venus Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data paper: https://
huggingface.co/papers/2604.19
859
…
RESEARCH
-

DR-Venus: Frontier Edge-Scale Deep Research Agents with 10K Data
By
–
-

Near-Future Policy Optimization Research Paper
By
–
Near-Future Policy Optimization paper: https://
huggingface.co/papers/2604.20
733
… -

RLVR Effectiveness in Low Data Compute Regimes Study
By
–
Our MLSys 2026 paper is live on arXiv: “Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes.” @realjustinbauer @Walshe_tech @pham_derek @harit_v @ArminPCM @fredsala and @paroma_varma present a comprehensive empirical study of open-source SLMs
-
IEEE Magazine Publication Paradox: Article From Future
By
–
It's 2026. I received a copy of IEEE BITS magazine in the mail. It contains my robustness article that was accepted in 2025. The magazine date is September 2024, before I finished writing the article Separately, my toddler was excited to get a "book" in the mail ft. his dad
-
Fraudulent Author Names Discovered in Academic Paper
By
–
all the author names were fraudulent; they had nothing to do with the paper
-
AI Discovery of MRSA-Effective Pill Could Transform Treatment
By
–
If AI can find a pill effective vs MRSA that would be big.
-

MIT Improves Reasoning Model Confidence Calibration Through RL Training
By
–
How do top reasoning models become overconfident? MIT found that RL rewards correct answers w/o considering how sure the model is. By training them to estimate their confidence about each answer, the team boosted uncertainty estimates w/o hurting accuracy:
-
Decoupled DiLoCo Training System Enables Resilient Large Scale AI
By
–
It's been a delight to provide small amounts of advice and suggestions to people working on the Decoupled DiLoCo training system. This approach enables graceful handling of failures in large scale training jobs, by allowing (N-1) / N units to proceed when one fails.
— Jeff Dean (@JeffDean) 23 avril 2026
Thread ⬇️ https://t.co/z97PgtNBuuIt's been a delight to provide small amounts of advice and suggestions to people working on the Decoupled DiLoCo training system. This approach enables graceful handling of failures in large scale training jobs, by allowing (N-1) / N units to proceed when one fails. Thread
-
Claude GPT agents robots breakthrough medical advances
By
–
Ces derniers jours donnent le vertige : Claude et GPT franchissent un nouveau cap, ouvrant un champ des possibles sans précédent. Google lance ses agents IA. La Chine fait courir un semi-marathon à ses robots humanoïdes. Des maladies jusqu'ici redoutables sont sur le point