AI Dynamics

Global AI News Aggregator

About

@jiqizhixin

  • VLMgineer: AI-Powered Robots Design Their Own Tools
    VLMgineer: AI-Powered Robots Design Their Own Tools

    Can AI truly empower robots to invent their own solutions? George Jiayuan Gao, Tianyu Li, and colleagues from UPenn present VLMgineer. This framework leverages Vision Language Models (VLMs) to brainstorm initial tool designs and action plans. It then refines these ideas using evolutionary search in simulation, optimizing both the tool's geometry and how the robot uses it. VLMgineer consistently outperforms existing human-crafted tools and VLM-generated designs from human specifications across diverse, challenging everyday manipulation tasks, transforming complex robotics problems into straightforward executions. VLMgineer: Vision Language Models as Robotic Toolsmiths Project: vlmgineer.github.io Paper: arxiv.org/abs/2507.12644 Our report: mp.weixin.qq.com/s/FXdeQhAeq… 📬 #PapersAccepted by Jiqizhixin

    → View original post on X — @jiqizhixin, 2026-04-03 14:45 UTC

  • HACRL: AI Agents Learn Together Without Losing Autonomy
    HACRL: AI Agents Learn Together Without Losing Autonomy

    What if diverse AI agents could mutually learn and improve without sacrificing their autonomy? Researchers from Beihang University, Bytedance China, Tsinghua University, and Peking University have just unveiled Heterogeneous Agent Collaborative Reinforcement Learning (HACRL)! This innovative framework allows different types of AI agents to share verified learning experiences during training, creating a bidirectional flow of knowledge to enhance performance for everyone. Unlike other multi-agent systems, it requires no coordinated deployment and fosters true peer-to-peer growth, not one-way teaching. Their HACPO algorithm consistently boosts all participating agents, outperforming GSPO by 3.3% on diverse reasoning benchmarks while dramatically cutting training data costs in half. Heterogeneous Agent Collaborative Reinforcement Learning Paper: arxiv.org/abs/2603.02604 Github Page: zzx-peter.github.io/hacrl/ Huggingface: huggingface.co/papers/2603.0… Our report: mp.weixin.qq.com/s/ggzim_4Pc… 📬 #PapersAccepted by Jiqizhixin

    → View original post on X — @jiqizhixin, 2026-04-03 12:43 UTC

  • Streamo: Real-time Streaming Video LLM for Intelligent Assistance
    Streamo: Real-time Streaming Video LLM for Intelligent Assistance

    What if an AI could truly understand live video streams and act as your intelligent assistant, in real-time? Researchers from Hong Kong Baptist University and Tencent Youtu Lab just unveiled a major step forward! They present Streamo, a real-time streaming video LLM. It's trained on a new, massive instruction dataset (Streamo-Instruct-465K) to enable unified understanding across many streaming video tasks. Streamo excels at real-time narration, complex action understanding, event captioning, and time-sensitive Q&A. It bridges the gap between static video analysis and genuinely interactive, intelligent multimodal AI assistants in continuous streams! Streaming Instruction Tuning Project: jiaerxia.github.io/Streamo/ Code: github.com/maifoundations/St… Our report: mp.weixin.qq.com/s/Q28azqwk-… 📬 #PapersAccepted by Jiqizhixin

    → View original post on X — @jiqizhixin, 2026-04-03 03:36 UTC

  • Anthropic Research: Emotion Concepts in Large Language Models
    Anthropic Research: Emotion Concepts in Large Language Models

    Interesting 🤔 Anthropic (@AnthropicAI) New Anthropic research: Emotion concepts and their function in a large language model. All LLMs sometimes act like they have emotions. But why? We found internal representations of emotion concepts that can drive Claude’s behavior, sometimes in surprising ways. — https://nitter.net/AnthropicAI/status/2039749628737019925#m

    → View original post on X — @jiqizhixin, 2026-04-03 01:30 UTC

  • Action-to-Action Flow Matching: Ultra-Fast Robot Control Method
    Action-to-Action Flow Matching: Ultra-Fast Robot Control Method

    What if real-time robot control didn't have to wait for slow, iterative action generation? MARS Lab at Nanyang Technological University (Jindou Jia et al.) introduces Action-to-Action Flow Matching (A2A). This novel method uses a robot's own historical actions to directly predict the next move, skipping the slow, random noise sampling of traditional diffusion models. A2A enables lightning-fast, single-step action generation (0.56 ms!), vastly outperforming existing methods in speed, training efficiency, robustness to visual noise, and generalization to unseen configurations. It even shows versatility in video generation! Action-to-Action Flow Matching Website: lorenzo-0-0.github.io/A2A_Fl…  arXiv: arxiv.org/pdf/2602.07322  Code: github.com/JIAjindou/A2A_Flo… Our report: mp.weixin.qq.com/s/mrSUcVLUA… 📬 #PapersAccepted by Jiqizhixin

    → View original post on X — @jiqizhixin, 2026-04-02 18:30 UTC

  • Memory Sparse Attention Framework Enables 100M Token Processing
    Memory Sparse Attention Framework Enables 100M Token Processing

    Can AI models finally process context the size of a lifetime? Evermind, Shanda Group, and Peking University present Memory Sparse Attention (MSA)! This new framework gives AI a massively scalable, end-to-end trainable long-term memory. It uses an innovative sparse attention architecture and other techniques to handle hundreds of millions of tokens with linear efficiency, maintaining exceptional precision. MSA processes 100M tokens on 2xA800 GPUs with less than 9% precision degradation from 16K. It significantly outperforms frontier LLMs, SOTA RAG systems, and leading memory agents in long-context benchmarks, paving the way for lifetime-scale AI memory. MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens Code: github.com/EverMind-AI/MSA Paper: zenodo.org/records/19103670 Our report: mp.weixin.qq.com/s/FHJA4kALc… 📬 #PapersAccepted by Jiqizhixin

    → View original post on X — @jiqizhixin, 2026-04-02 14:25 UTC

  • LeWorldModel: Stable End-to-End JEPA from Pixels
    LeWorldModel: Stable End-to-End JEPA from Pixels

    LeWorldModel: Stable End-to-End JEPA from Pixels Paper: https://
    le-wm.github.io
    Project: https://
    arxiv.org/pdf/2603.19312
    v1
    … Our report: https://
    mp.weixin.qq.com/s/VycD8SODNAnH
    iLIg4DlZvw

    → View original post on X — @jiqizhixin

  • EmoStyle: AI Framework Transforms Images to Evoke Human Emotions
    EmoStyle: AI Framework Transforms Images to Evoke Human Emotions

    What if AI could stylize images to truly evoke specific human emotions? Jingyuan Yang, Zihuan Bai, and Hui Huang from CSSE, Shenzhen University They introduce EmoStyle, a groundbreaking framework that transforms your images to reflect emotions like 'amusement' or 'disgust'

    → View original post on X — @jiqizhixin

  • NS-Diff: AI Framework for Realistic Fluid Video Generation
    NS-Diff: AI Framework for Realistic Fluid Video Generation

    Can AI finally generate videos that look and feel real, especially with fluids? Researchers Zijun Deng and Yuxin Peng from Peking University unveil NS-Diff. Their new framework intelligently identifies fluid and rigid elements within noisy video frames. It then injects

    → View original post on X — @jiqizhixin

  • AI Breakthrough in Microscale Simulation
    AI Breakthrough in Microscale Simulation

    Can AI finally unlock the secrets of life's smallest structures through simulation? The Chinese University of Hong Kong, Shenzhen and collaborators present a groundbreaking AI. They first show existing video generation models fail at accurately simulating microscale phenomena.

    → View original post on X — @jiqizhixin