What if you could tell an AI to generate audio starting at exactly 3 seconds, with perfect speech clarity? Researchers from Tsinghua University and Shengshu AI (with USTC and Monash) present ControlAudio. Instead of just typing a prompt, this system handles three instructions
@jiqizhixin
-

Monet-7B: AI Reasoning Directly in Visual Latent Space
By
–
What if your AI could truly “think in pictures,” not just describe them? Researchers from Peking University, Kling Team, and Amazon AGI introduce Monet-7B, a new framework that lets multimodal AI reason directly inside visual latent space—no external tools needed. Instead of
-

Strategy Genes: AI Evolution Through Compact Reusable Code
By
–
Can you turn messy trial-and-error into pure strategy, like a game AI evolving mid-battle? From Tsinghua University and EvoMap, researchers challenge how AI reuses past experience. Instead of bulky "skill manuals," they propose "Strategy Genes" — compact, evolution-ready code
-

Beyond Model Size: Smart Scaffolding for LLM Agents
By
–
Think building an AI agent is just a better brain? What if the real secret is what you add outside the model? A team from Shanghai Jiao Tong University, OPPO, and others argues that the future of LLM agents isn't about bigger weights, but smarter scaffolding. They introduce a
-

Unified Framework Controls Large Language Models Without Breaking
By
–
Can you control a large language model without breaking its brain? Zhejiang University and Alibaba Group researchers just showed how. They unify all model control methods (fine-tuning, LoRA, activation edits) into one single framework, separating effects into "preference"
-

Huawei AURA: Real-Time AI Video Processing System
By
–
What if your AI could watch a live video feed and answer your questions, or even alert you, all in real time? Researchers from Huawei Research and CUHK MMLab, introduce AURA. This new system allows a single video AI to continuously process live streams, enabling instant
-

Single Photo to Rigged 3D Character Generation with AniGen
By
–
What if you could generate a fully rigged, animatable 3D character from just a single photo? Researchers from The University of Hong Kong, VAST, CUHK, and Tsinghua University present AniGen for exactly that. They created a unified system that simultaneously generates a 3D
-

Security research evaluates Claude Code’s auto mode risk management
By
–
How secure is the new "auto mode" for AI coding assistants? Researchers from HKUST and ETH Zurich put Claude Code's permission system to the test. They found it misses dangerous actions when a task's risk is ambiguous. In a stress test, it missed 81% of risky actions, far
-

QuatRoPE: A New Method for 3D Spatial Encoding in LLMs
By
–
How can we teach AI to truly understand 3D space? Researchers from Peking University & SUSTech present QuatRoPE. It's a new method that efficiently encodes the 3D relationships between objects for Large Language Models, using a scalable linear approach instead of a messy
-

DeepSeek-V4 Research Paper Released Focusing on Million-Token Context Efficiency
By
–
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence Paper: https://
huggingface.co/deepseek-ai/De
epSeek-V4-Pro/blob/main/DeepSeek_V4.pdf
…
