AI Dynamics

Global AI News Aggregator

About

@aiatmeta

  • Reinforcement Learning Optimizes Model Reasoning with Token Efficiency

    RL trains our models to "think" before they answer, a process known as test-time reasoning. To serve this capability to billions of users and efficiently use tokens, we rely on two key levers: thinking time penalties to optimize token use and multi-agent orchestration that boosts

    → View original post on X — @aiatmeta

  • Scaling Properties of Muse Spark: Pretraining, RL, and Reasoning
    Scaling Properties of Muse Spark: Pretraining, RL, and Reasoning

    To build personal superintelligence, our model’s capabilities should scale predictably and efficiently. Below, we share how we study and track Muse Spark’s scaling properties along three axes: pretraining, reinforcement learning, and test-time reasoning. Let’s start with

    → View original post on X — @aiatmeta

  • Meta’s Muse Spark Returns Company to Frontier AI Race
    Meta’s Muse Spark Returns Company to Frontier AI Race

    Meta is back! Muse Spark scores 52 on the Artificial Analysis Intelligence Index, behind only Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. Muse Spark is the first new release since Llama 4 in April 2025 and also Meta's first release that is not open weights Muse Spark is a new model from @Meta evaluated on Artificial Analysis. We were given early access by Meta to independently benchmark the model. It is the first frontier-class model from Meta since Llama 4 Maverick was released in April 2025, and notably the first @AIatMeta model that is not being released as open weights. The release follows Meta's reorganization of its AI efforts under Meta Superintelligence Labs, and signals that Meta is re-entering the frontier race after roughly a year of relative quiet. For context, Llama 4 Maverick and Scout scored 18 and 13 respectively on the Artificial Analysis Intelligence Index as non-reasoning models at the time of their release, while Muse Spark scores 52. Muse Spark essentially closes the gap between to the frontier in a single release. The model is not open source and is not yet accessible via an API but Meta has shared they expect this to come soon. Meta is also integrating Muse Spark into their first party products including their Meta AI chat product, Facebook, Instagram and Threads. Key takeaways from our benchmarks: ➤ Muse Spark scores 52 on the Artificial Analysis Intelligence Index, placing it within the top 5 models we have benchmarked. It sits ahead of Claude Sonnet 4.6, GLM-5.1, MiniMax-M2.7, Grok 4.20 and behind Gemini 3.1 Pro Preview, GPT-5.4 and Claude Opus 4.6 ➤ Muse Spark is notably token efficient for its intelligence level. It used 58M output tokens to run the Intelligence Index, comparable to Gemini 3.1 Pro Preview (57M) and notably lower than Claude Opus 4.6 (Adaptive Reasoning, max effort, 157M), GPT-5.4 (xhigh, 120M) and GLM-5 (110M) ➤ Muse Spark is the second-most capable vision model we have benchmarked. It scores 80.5% on MMMU-Pro, behind only Gemini 3.1 Pro Preview (82.4%) ➤ Muse Spark performs strongly on reasoning and instruction-following evaluations. It scores 39.9% on HLE, trailing only Gemini 3.1 Pro Preview (44.7%) and GPT-5.4 (xhigh, 41.6%). The model also achieved 5th highest in CritPT with a score of 11%, an eval that is focused on difficult physics research questions. This is substantially above above Gemini 3 Flash (9%) and Claude 4.6 Sonnet (3%) ➤ Agentic performance does not stand out. On GDPval-AA, our evalaution focused on real world work tasks, Muse Spark scores 1427, behind both Claude Sonnet 4.6 at 1648 and GPT-5.4 at 1676, but ahead of Gemini 3.1 Pro Preview at 1320. On On TerminalBench Hard, Muse Spark trails Claude Sonnet 4.6, GPT-5.4, and Gemini 3.1 Pro. Muse Spark joins others in achieving a high τ²-Bench Telecom score of 92% Key model details: ➤ Modalities: Multimodal including text and vision input, text output ➤ License: Proprietary, Meta's first frontier model not released as open weights ➤ Availability: No public API at the time of publishing. Meta expects to provide API access soon. Meta has started integration into their first party AI offering Meta AI and inside Facebook, Instagram, and Threads

    → View original post on X — @aiatmeta, 2026-04-08 16:16 UTC

  • Muse Spark: Advanced Visual AI for Interactive STEM and Home Solutions

    Muse Spark is built from the ground up to integrate visual information across domains and tools. It achieves strong performance on visual STEM questions, entity recognition, and localization, enabling interactive experiences like troubleshooting your home appliances with dynamic

    → View original post on X — @aiatmeta

  • Personal Superintelligence for Health Education with Physician-Curated Data

    Personal superintelligence will help people learn about their health. We collaborated with 1,000+ physicians to curate training data that enables more factual and comprehensive responses. It can generate interactive displays that unpack and explain health information such as the

    → View original post on X — @aiatmeta

  • Muse Spark: First Step in AI Scaling Ladder
    Muse Spark: First Step in AI Scaling Ladder

    Muse Spark is the first step on our scaling ladder and the first product of a ground-up overhaul of our AI efforts. It offers competitive performance in multimodal perception, reasoning, health, and agentic tasks. We continue to invest in areas with current performance gaps,

    → View original post on X — @aiatmeta

  • Muse Spark Launches Contemplating Mode for Advanced Reasoning
    Muse Spark Launches Contemplating Mode for Advanced Reasoning

    We're also releasing Contemplating mode, which orchestrates multiple agents that reason in parallel. This allows Muse Spark to compete with the extreme reasoning modes of frontier models such as Gemini Deep Think and GPT Pro. Contemplating will be rolling out gradually in

    → View original post on X — @aiatmeta

  • SAM 3.1 Object Multiplexing: Track 16 Objects Simultaneously
    SAM 3.1 Object Multiplexing: Track 16 Objects Simultaneously

    The core innovation in SAM 3.1 is object multiplexing, allowing the model to track up to 16 objects in a single forward pass. Previously, each object required its own dedicated pass, but with multiplexing, SAM 3.1 processes all tracked objects together, eliminating redundant

    → View original post on X — @aiatmeta

  • SAM 3.1: Object Multiplexing Enhances Video Processing Efficiency
    SAM 3.1: Object Multiplexing Enhances Video Processing Efficiency

    We’re releasing SAM 3.1: a drop-in update to SAM 3 that introduces object multiplexing to significantly improve video processing efficiency without sacrificing accuracy. We’re sharing this update with the community to help make high-performance applications feasible on smaller,

    → View original post on X — @aiatmeta

  • TRIBE v2 Predicts Brain Responses Without Retraining
    TRIBE v2 Predicts Brain Responses Without Retraining

    Without any retraining, TRIBE v2 can reliably predict the brain responses of individuals it has never seen before, achieving a nearly 2-3x improvement over previous methods for both movies and audiobooks We’re releasing the model, codebase, paper, and demo to help researchers

    → View original post on X — @aiatmeta