AI Dynamics

Global AI News Aggregator

About

AGENTS

  • Keynote-Worthy Call for Thoughtfulness Amid AI Agent Hype

    this is keynote worthy and means even more coming from a coding agent creator. i am always pro thoughtfulness and we love featuring calls for sanity amongst the hype (my agents conf opened with @sayashk pointing out how most agents are overhyped and it was so well received)

    → View original post on X — @swyx

  • ARC-AGI-3 Benchmarks Agent Performance Against Human Action Efficiency
    ARC-AGI-3 Benchmarks Agent Performance Against Human Action Efficiency

    ARC-AGI-3 scores agents on how close they are to human action efficiency. All ARC-AGI-3 environments were solved by at least 2 human testers out of 10 (most of the time it was 5+). We use the action count of the 2nd best tester (to avoid outlier performance) as our human

    → View original post on X — @fchollet

  • The AI Scientist Published in Nature: Fully Automated Research
    The AI Scientist Published in Nature: Fully Automated Research

    The AI Scientist: Towards Fully Automated AI Research, Now Published in Nature!!✨ Today in Nature we share a comprehensive technical summary of our work on The AI Scientist, including new scaling law results showing how it improves with more compute and more intelligent foundation models. The AI Scientist autonomously creates its own research ideas, codes up and conducts experiments to test those ideas, creates figures to visualize the results, writes an entire scientific manuscript summarizing what it has discovered, and conducts its own “peer” review of the resulting paper. One of its papers–entirely AI generated–passed peer review at a top-tier AI conference workshop, a historic milestone marking the dawn of a new era of AI-accelerated scientific discovery. 🔬🧪✨🧬💡🔭 Paper nature.com/articles/s41586-0… Blog sakana.ai/ai-scientist-natur… Work done in collaboration with a great team from Sakana, Oxford, and my lab at UBC. Thanks and congratulations everyone! @_chris_lu_ @cong_ml @RobertTLange @_yutaroyamada @shengranhu @j_foerst @hardmaru

    → View original post on X — @_yutaroyamada, 2026-03-25 18:01 UTC

  • ARC-AGI-3 Benchmark Monitors Frontier Models AGI Breakthrough

    At the moment, ARC-AGI-3 is the only unsaturated agentic AI benchmark. Sub-1% scores from frontier models on the private test set. If you want to be among the first to know when an AGI breakthrough happens, monitor the ARC-AGI-3 leaderboard. Any sudden score jump will mean

    → View original post on X — @fchollet

  • ARC-AGI-2 Kaggle Competition Final Round With Unlimited Prize

    We're also running one last ARC-AGI-2 competition on Kaggle this year. Get your high score in: since this is the last official ARC-AGI-2 competition, the grand prize will go to the top score regardless of whether it's above the 85% threshold.

    → View original post on X — @fchollet

  • ARC-AGI-3 Kaggle Competition Tests AI Agents

    You can also enter the ARC-AGI-3 competition on Kaggle. Your AI agents will be tested on two separate private test sets of 55 environments.

    → View original post on X — @fchollet

  • ARC-AGI-3 Benchmark Evaluates Agentic Intelligence Systems

    ARC-AGI-3 is out now! We've designed the benchmark to evaluate agentic intelligence via interactive reasoning environments. Beating ARC-AGI-3 will be achieved when an AI system matches or exceeds human-level action efficiency on all environments, upon seeing them for the first

    → View original post on X — @fchollet

  • ARC-AGI-3: New Benchmark Shows AI Lacks True Learning Ability

    Announcing ARC-AGI-3 The only unsaturated agentic intelligence benchmark in the world Humans score 100%, AI <1% This human-AI gap demonstrates we do not yet have AGI Most benchmarks test what models already know, ARC-AGI-3 tests how they learn

    → View original post on X — @lmthang, 2026-03-25 17:37 UTC

  • The AI Scientist Published in Nature: Automated Scientific Discovery Milestone
    The AI Scientist Published in Nature: Automated Scientific Discovery Milestone

    I am really excited to share that our work on The AI Scientist has been published in Nature Automated Scientific Discovery has been something I only dreamt about at the start of my PhD. Today, we are making big leaps into a world in which autonomous agents support human researchers in tackling some of the most fundamental problems. In August 2024, The AI Scientist-v1 showed first sparks of LLM agents becoming capable of conducting research end-to-end. While the generated artifacts were still far from perfect, it was clear that automated discovery was about to change. We scaled the system and improved all ingredients of the pipeline. In April 2025, The AI Scientist-v2 had become capable of producing a paper that could pass the human peer review of an ICLR workshop. This is only the beginning. Systems like AlphaEvolve, ShinkaEvolve, AIDE, and Autoresearch will continue to shape the future of how research is conducted. Our METR-style scaling results indicate that model improvements have direct downstream impacts. Still, there are many challenges. Both technical and societal. I have a strong belief that we, as a collective, will find the answers and adapt. This has been an enormous amount of work by an outstanding set of human researchers @_chris_lu_ @cong_ml @_yutaroyamada @shengranhu @j_foerst @jeffclune @hardmaru @SakanaAILabs with many long nights of work. I am super grateful for the entire ride, learnings and the future to come. Thank you to everyone! Sakana AI (@SakanaAILabs) The AI Scientist: Towards Fully Automated AI Research, Now Published in Nature Nature: nature.com/articles/s41586-0… Blog: sakana.ai/ai-scientist-natur… When we first introduced The AI Scientist, we shared an ambitious vision of an agent powered by foundation models capable of executing the entire machine learning research lifecycle. From inventing ideas and writing code to executing experiments and drafting the manuscript, the system demonstrated that end-to-end automation of the scientific process is possible. Soon after, we shared a historic update: the improved AI Scientist-v2 produced the first fully AI-generated paper to pass a rigorous human peer-review process. Today, we are happy to announce that “The AI Scientist: Towards Fully Automated AI Research,” our paper describing all of this work, along with fresh new insights, has been published in @Nature! This Nature publication consolidates these milestones and details the underlying foundation model orchestration. It also introduces our Automated Reviewer, which matches human review judgments and actually exceeds standard inter-human agreement. Crucially, by using this reviewer to grade papers generated by different foundation models, we discovered a clear scaling law of science. As the underlying foundation models improve, the quality of the generated scientific papers increases correspondingly. This implies that as compute costs decrease and model capabilities continue to exponentially increase, future versions of The AI Scientist will be substantially more capable. Building upon our previous open-source releases (github.com/SakanaAI/AI-Scien…), this open-access Nature publication comprehensively details our system's architecture, outlines several new scaling results, and discusses the promise and challenges of AI-generated science. This substantial milestone is the result of a close and fruitful collaboration between researchers at Sakana AI, the University of British Columbia (UBC) and the Vector Institute, and the University of Oxford. Congrats to the team! @_chris_lu_ @cong_ml @RobertTLange @_yutaroyamada @shengranhu @j_foerst @hardmaru @jeffclune — https://nitter.net/SakanaAILabs/status/2036840833690071450#m

    → View original post on X — @_yutaroyamada, 2026-03-25 17:31 UTC

  • Fleet introduces shareable skills for team domain knowledge

    Fleet now has shareable skills. Capture your team's domain knowledge once, attach it to any agent, and share it across your workspace. Create skills from a prompt or previous chat, write them manually, or use a template. Read more: https://
    blog.langchain.com/skills-in-lang
    smith-fleet/?utm_medium=social&utm_source=twitter&utm_campaign=q1-2026_fleet-launch_aw

    → View original post on X — @langchain