AI Dynamics

Global AI News Aggregator

About

MACHINE LEARNING

  • Meta Harnesses: Automated Framework Optimization for AI Tasks
    Meta Harnesses: Automated Framework Optimization for AI Tasks

    Meta Harnesses is Autoresearch on steroids. Something I've been exploring recently is to get long running agents to hill climb on a verifiable task to continuously improve without my intervention. Karpathy's Autoresearch did this pretty well on specific tasks, but this weekend I tried Meta Harnesses which moves one level of abstraction up. What does Meta Harness do? Autoresearch can be used in harness like Claude Code / Codex to generate experiments to try, evaluate results, and continue looping. Meta Harness generates a harness itself that optimizes on a task or a set of task. Here, we define a harness as "a single-file Python program that modifies task-specific prompting, retrieval, memory, and orchestration logic". The idea is that LLMs are very powerful today, but to harness [pun intended] their power, you need to give it the right prompts and context. Meta Harnesses automates coming up with the right prompts and the right way to retrieve context to solve a problem. Where did this idea come from? This is from a paper from Stanford and the author of DSPy written last week. The paper shows fantastic performance on 3 tasks: text classification, math reasoning (IMO level problems) and coding (Terminal Bench 2.0), far outperforming traditional harnesses. The discovered harnesses are interesting: math for example, splits up the logic into different categories (Combinatorics, Geometry, Number Theory, Algebra) and prompts and looks at the context differently. The coding harness, amongst other things, pre-processes the tools available in the environment to save exploratory turns. When should you use and not use it? Meta Harnesses seem pretty useful for tackling a specific but wide set of problems where the result is verifiable. In contrast, when I tried it on a specific task like Chess, it arbitrarily divides the problem into separate tasks – opening, mid game, end game, and creates different approaches for each. This "works" but isn't really clean because we believe there should be one approach that does all three. It does far better on things like examinations (JEE, Gaokao) where it splits problems into categories and tackles each category with different strategies. This paper covers a pretty light version of what a harness means. In the future, we can split up tasks into harnesses that have access to specific kinds of data, specific toolchains and various models to get even better results. Overall, pretty cool applied AI approach to hillclimb a verifiable task in a specific domain with variety within the problem space.

    → View original post on X — @askalphaxiv, 2026-04-06 16:22 UTC

  • Statistics and Machine Learning: Tom Mitchell interviews Michael I. Jordan

    New podcast episode for you! The field of Statistics enters the Machine Learning conversation, as @tommmitchell talks with Michael I. Jordan of @inria_paris! Watch here: piped.video/BfBxS2rEH3k

    → View original post on X — @stanfordhai, 2026-04-06 16:18 UTC

  • SF Crime Progress with Flock Safety Technology

    RT @aleximm: We can just do things. Like solve crime. I'm thrilled by the progress SF has made, and I'm proud Flock Safety has played an i…

    → View original post on X — @scobleizer, 2026-04-06 15:47 UTC

  • Audit reveals Claude Code’s token waste

    This user audited 926 sessions of Claude Code and discovered that most of the token waste came from its platform. Everyone blames Anthropic for the limitations, so he decided to analyze the data. 858 sessions, 18,903 turns, and an estimated spend of $1,619 in 33 days. This is

    → View original post on X — @s0n_ia_

  • MIT CSAIL Introduces OSGym for Training Computer Agents
    MIT CSAIL Introduces OSGym for Training Computer Agents

    How do you train AI agents that can use computers like humans? 🧵 MIT CSAIL researchers introduce "OSGym": scalable OS infrastructure to improve the capabilities of computer use agents. It introduces large-scale training made possible by extensive infrastructure optimization: bit.ly/47JprPd [Translated from EN to English]

    → View original post on X — @mit_csail, 2026-04-06 15:25 UTC

  • Google’s AI Predicts Flash Floods 24 Hours in Advance

    Google's new AI can predict flash floods 24 hours before they strike. How it works: > Uses Gemini to extract confirmed flood locations and times from global news > Builds a dataset of past events that never formally existed. > That dataset feeds a neural network > The neural network combines real-time weather forecasts with local terrain, soil absorption rates, and urban density > Can flag at-risk areas with results that match the accuracy of America's National Weather Service > Can do this in countries that have almost no flood monitoring infrastructure at all Flash floods kill more than 5,000 people every year, but most victims never see them coming. A 12-hour warning alone can reduce flash flood damage by 60%. …For billions of people without any early warning system, this could be the difference between life and death. [Translated from EN to English]

    → View original post on X — @rowancheung, 2026-04-06 15:13 UTC

  • Andrej Karpathy’s LLM Wiki: Persistent Memory vs Traditional RAG
    Andrej Karpathy’s LLM Wiki: Persistent Memory vs Traditional RAG

    🚨 Andrej Karpathy just dropped something that could replace a lot of RAG workflows. It's called LLM Wiki. The idea is simple: Most AI systems retrieve context from scratch every time you ask a question. LLM Wiki doesn't. It builds a persistent knowledge base that gets better every time you add a new source. So instead of: • search docs
    • pull fragments
    • answer
    • forget everything
    • repeat it does this: • ingest a source
    • extract the important ideas
    • update entity pages
    • revise topic summaries
    • connect related concepts
    • flag contradictions
    • keep compounding the knowledge over time That shift matters. RAG is great for retrieval. But a lot of people are really trying to build memory. Not just "find me the right chunk again."
    More like: "help me build an evolving model of this topic over time." That's what this is. Karpathy's examples are strong too: • personal knowledge
    • long-horizon research
    • books and topics
    • internal company knowledge
    • meeting transcripts
    • customer calls Basically, anything where the knowledge should accumulate, not reset every session. The best way to think about it: Obsidian is the IDE.
    The LLM is the programmer.
    The wiki is the codebase. You don't manually maintain the system. You feed it sources, ask questions, and the AI keeps the structure alive. That's a much bigger idea than "better RAG." 100% open source. [Translated from EN to English]

    → View original post on X — @scobleizer, 2026-04-06 15:06 UTC

  • New Scaling Laws for 350M Model Training Tokens
    New Scaling Laws for 350M Model Training Tokens

    FACT: If you don't train your 350M model on 28T tokens, you're not optimal Nicholas Roberts (@nick11roberts) That new LFM2.5-350M is super overtrained, right? And everyone was shocked about how far they pushed it? As it turns out, we have a brand new scaling law for that! 🧵 [1/n] — https://nitter.net/nick11roberts/status/2041141606305124486#m

    → View original post on X — @maximelabonne, 2026-04-06 15:05 UTC

  • OpenSeeker: AI-Native Search Beyond Keyword Matching

    OpenSeeker: Rethinking Search With AI-Native Reasoning In this episode of Artificial Intelligence: Papers and Concepts, we explore OpenSeeker, an emerging approach to building AI-native search systems that go beyond traditional keyword matching. Instead of retrieving links based purely on queries, OpenSeeker focuses on reasoning over information helping users get structured, context-aware answers rather than a list of results. We break down how modern search is evolving with large language models, why retrieval alone is no longer enough, and how systems like OpenSeeker combine retrieval with reasoning to deliver more accurate and useful outputs. If you’re interested in AI-powered search, retrieval-augmented generation, or the future of information discovery, this episode explains why OpenSeeker represents a shift toward more intelligent and answer-driven search experiences. Resources: Paper Link: arxiv.org/abs/2603.15594v1 Interested in Computer Vision and AI consulting and product development services? Email us at contact@bigvision.ai or visit us at bigvision.ai

    → View original post on X — @learnopencv, 2026-04-06 14:30 UTC

  • RAG, AI Agent, Fine-Tuning, and LLM Customization Strategy Explained
    RAG, AI Agent, Fine-Tuning, and LLM Customization Strategy Explained

    RAG, AI Agent, Fine-Tuning, LLM Customization Strategy Briefly Explained! #BigData #Analytics #DataScience #AI #MachineLearning #NLProc #LLM #IoT #IIoT #PyTorch #Python #RStats #TensorFlow #Java #JavaScript #ReactJS #GoLang #CloudComputing #Serverless #DataScientist #Linux #Programming #Coding #100DaysofCode geni.us/RAG-AI-Agent

    → View original post on X — @gp_pulipaka, 2026-04-06 14:26 UTC