It's surprising to see the pace of progress of LLMs (which until not long ago were very bad at math) achieving ever greater milestones. Take a look at the graph above. The progress curve is getting faster and faster. I'll leave the link to the paper here: https://
arxiv.org/abs/2605.06651
GENERATIVE AI
-
LLMs’ Math Progress Accelerates Rapidly
By
–
-

MARBLE: Multi-Aspect Reward Balance for Diffusion RL
By
–
MARBLE Multi-Aspect Reward Balance for Diffusion RL paper: https://
huggingface.co/papers/2605.06
507
… -

Google reclaims lead in FrontierMath Tier 4
By
–
LEADER in FRONTIER MATH T4! New record in one of the most challenging math benchmarks -FrontierMath Tier 4- where Google has just reclaimed the lead with a 47.9% !!, stealing the position from GPT 5.5 Pro It does it with its new math agent
-

New Continuous Latent Diffusion Language Model paper
By
–
Continuous Latent Diffusion Language Model paper: https://
huggingface.co/papers/2605.06
548
… -

Using Claude as a personal multimedia system with custom prompts
By
–
J’ai arrêté de payer Spotify. Résilié Netflix. Résilié Disney+. Maintenant, mon laptop est devenu un meilleur centre de divertissement…gratuitement. Grâce à Claude. Voici 8 prompts qui ont transformé l’IA en système multimédia personnel : [ Ajoutez en signet pour ne pas
-

HuggingFace releases autonomous ML engineer
By
–
HuggingFace released ml-intern, an autonomous ML engineer! ml-intern is an agent that reads papers, writes ML code, trains models, and ships production-ready models using the Hugging Face ecosystem. The workflow: You give it a task like "fine-tune Llama on my dataset." It
-
Claude Dreams and Continuously Improves
By
–
— AI Claude no longer sleeps—it dreams. Claude will now operate 24/7. (A prerequisite for AGI? Yes.) Anthropic has just launched 'Dreams': between conversations, your AI agent reviews its past sessions and improves on its own. The next day, it’s better than yesterday.
-
Technical Workshop on Deploying and Managing AI Agents
By
–
In this meet-up, we'll cover how to: Deploy agents with durable execution so runs survive crashes, deploys, and long waits for human input Use checkpoints and memory stores Add human-in-the-loop gates for consequential decisions We'll also discuss how teams
-
User switches from Grok 4.1 Fast to Grok 4.3 without reasoning
By
–
I just use whatever is available, yday that was Grok 4.1 Fast Non Reasoning, now I try Grok 4.3 via grok-latest with reasoning set to none
-

RAG is old way; future AI memory is compilation, not retrieval
By
–
RAG is already becoming the “old way” The future of AI memory is not retrieval.
It’s compilation. Here’s the shift in one sentence: From searching information To structuring knowledge The new model? LLM Wiki Instead of: Chunking documents Running similarity