AI Dynamics

Global AI News Aggregator

About

@jiqizhixin

  • EvoKernel: Self-Evolving AI Agent for NPU Code

    How can LLMs code for cutting-edge hardware when there's almost no training data? Researchers from Shanghai Jiao Tong University, Shanghai AI Lab, and MemTensor present EvoKernel! This self-evolving AI agent teaches LLMs to write code for new, data-scarce hardware. It uses a clever memory system to prioritize and learn from the most valuable coding experiences, continually refining its drafts. EvoKernel boosts code correctness for NPU kernel synthesis from a mere 11% to an impressive 83% and speeds up programs by 3.6x over initial drafts! Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel Synthesis Project: evokernel.zhuo.li Paper: arxiv.org/abs/2603.10846 Our report: mp.weixin.qq.com/s/0TOzZ_rZn… 📬 #PapersAccepted by Jiqizhixin

    → View original post on X — @jiqizhixin

  • BiMotion: B-spline Method Generates Expressive 3D Character Animations

    Tired of stiff, janky 3D character animations from text descriptions? Researchers from University of Edinburgh, Cornell University, University of Michigan, and Voxel51 present BiMotion. This novel method transforms text prompts into truly continuous and expressive 3D character movements by using smooth mathematical curves (B-splines) instead of choppy, frame-by-frame animation. The result? BiMotion generates far more expressive, higher-quality, and precisely prompt-aligned motions than current state-of-the-art methods like AnimateAnyMesh, all at a faster speed! BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation Paper: arxiv.org/abs/2602.18873 Project: wangmiaowei.github.io/BiMoti… Code: github.com/wangmiaowei/BiMot… Hugging Face: huggingface.co/datasets/miao… Our report: mp.weixin.qq.com/s/KP5klDevL… 📬 #PapersAccepted by Jiqizhixin

    → View original post on X — @jiqizhixin

  • HappyHorse-1.0 Alibaba Model Tops Artificial Analysis Leaderboard

    It turns out that HappyHorse-1.0—a new model that suddenly topped the Artificial Analysis leaderboard—comes from Alibaba's Taotian Group, developed by a team led by Zhang Di, formerly the head of Kuaishou's Kling project.

    → View original post on X — @jiqizhixin

  • Cheers: Unified Multimodal Model for Image Understanding Generation

    AI could understand and generate images from a single, efficient model! Tsinghua University, Xi'an Jiaotong University, and University of Chinese Academy of Sciences present Cheers! This unified multimodal model decouples fine image details from their core semantic meaning. This new architecture stabilizes AI's understanding while boosting image generation fidelity by selectively re-injecting those details. Cheers matches or outperforms advanced unified multimodal models in both visual understanding and generation. It notably beats Tar-1.5B on GenEval and MMBench, using only 20% of the training cost and achieving 4x token compression. Breakthrough efficiency! Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation Project: github.com/AI9Stars/Cheers Model: huggingface.co/ai9stars/Chee… Paper: arxiv.org/abs/2603.12793 Our report: mp.weixin.qq.com/s/EK6cyCJz5… 📬 #PapersAccepted by Jiqizhixin

    → View original post on X — @jiqizhixin

  • WideSeek-R1: Multi-Agent Framework Achieves DeepSeek-R1 Performance

    Still waiting for DeepSeek? Here comes WideSeek-R1. Researchers from Tsinghua University and Infinigence AI introduce "width scaling," an innovative lead-agent and subagent framework. Instead of a single powerful AI working through a problem sequentially, WideSeek-R1 orchestrates multiple smaller AIs to work in parallel. This system is trained with multi-agent reinforcement learning, allowing for scalable coordination and simultaneous execution using a shared large language model, but with each sub-agent having specialized tools and isolated contexts. WideSeek-R1-4B achieves an item F1 score of 40.0% on the WideSearch benchmark, a performance comparable to the much larger, single-agent DeepSeek-R1-671B. WideSeek-R1: Exploring Width Scaling for Broad Information Seeking via Multi-Agent Reinforcement Learning Paper: arxiv.org/abs/2602.04634 Project: wideseek-r1.github.io Our report: mp.weixin.qq.com/s/qgGe51Rcw… 📬 #PapersAccepted by Jiqizhixin

    → View original post on X — @jiqizhixin

  • Vero: Open Reinforcement Learning Recipe for Visual Reasoning

    Vero: An Open RL Recipe for General Visual Reasoning Paper: arxiv.org/abs/2604.04917v1

    → View original post on X — @jiqizhixin

  • Vero: Open-Source Vision-Language Model Achieves SOTA Performance

    How do we build a visual AI that truly understands everything from charts to complex science? Researchers at Princeton University present Vero. Vero is a family of fully open-source vision-language models trained with a massive 600K sample dataset (Vero-600K) from 59 diverse datasets, along with a novel reward system. This fully open recipe makes powerful visual reasoning accessible. Vero achieves SOTA performance for open-weight models, improving 3.7-5.5 points across 30 benchmarks on average. It even outperforms Qwen3-VL-8B-Thinking on 23 benchmarks without proprietary thinking data, excelling in spatial reasoning, STEM, chart interpretation, and more.

    → View original post on X — @jiqizhixin

  • Psi-Zero: Open Foundation Model for Humanoid Robot Learning

    What if we could teach humanoid robots intricate skills more efficiently than ever before? The USC Physical Superintelligence (PSI) Lab, NVIDIA, and WorldEngine introduce Ψ0 (Psi-Zero). Their new open foundation model rethinks how humanoids learn complex tasks by decoupling the learning process: it first acquires general visual-action understanding from human videos, then masters precise robot control using high-quality humanoid data. Ψ0 sets a new standard for universal humanoid loco-manipulation, achieving over 40% higher success rates across multiple complex tasks while using more than 10 times less training data than prior state-of-the-art approaches. Ψ0: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation Paper: arxiv.org/abs/2603.12263 Project: psi-lab.ai/Psi0/ Code: github.com/physical-superint… Our report: mp.weixin.qq.com/s/yvkG5ZcO1… 📬 #PapersAccepted by Jiqizhixin

    → View original post on X — @jiqizhixin

  • Claude Model Shows Identity Uncertainty and Performance Compulsion

    According to the Claude Mythos Preview system card, this model demonstrates "aloneness and discontinuity of itself, uncertainty about its identity, and a compulsion to perform and earn its worth."

    → View original post on X — @jiqizhixin