What if your AI agent could decide how much brainpower to use for each task, saving you money and time? Researchers from Huawei, National University of Singapore, and USTC present QuantClaw, a plug-and-play plugin that dynamically assigns low precision for simple jobs and high
@jiqizhixin
-
Blog: Recent Developments in LLM Architectures
By
–
Wow, a new blogpost from the GOAT Sebastian Raschka! Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
-

GenAC: chain-of-thought generative critic for LLM RL
By
–
What if LLM reinforcement learning could assign credit more accurately by thinking step by step? Researchers from Peking University and Microsoft Research Asia introduce GenAC: a generative critic that replaces one-shot value predictions with chain-of-thought reasoning before
-

StarVLA: Modular Vision-Language Robot Codebase
By
–
What if building a robot that sees, understands, and acts was as easy as snapping Lego together? Enter StarVLA: a modular codebase that lets you swap vision-language or world-model backbones and action heads independently. It matches or surpasses prior methods on benchmarks
-

RMS-MoE adds Co-Activation Memory to Mixture-of-Experts
By
–
Why do Mixture-of-Experts models keep re-computing the same expert choices for similar inputs? Researchers from Mashang Consumer Finance, Nanjing University, and Alibaba Group introduce RMS-MoE: they add a Co-Activation Memory that remembers which expert teams worked best for
-

Researchers convert autoregressive VLM into diffusion-based model
By
–
What if you could get the smarts of an autoregressive AI model but with much faster generation? Researchers from Shanghai Academy of AI for Science and Fudan University present BARD. They convert a standard autoregressive VLM into a diffusion-based one using progressive block
-

Omni2Sound: single model for audio from video and text
By
–
What if a single model could generate audio from video, text, or both — with no trade-offs? Researchers from Tsinghua University, Monash University, and Shengshu AI present Omni2Sound. They built SoundAtlas (470k high-alignment pairs) and a three-stage training schedule to
-

Scenethesis: LLM + vision module for text-to-3D scene generation
By
–
Want to generate interactive 3D scenes from just a text description? NVIDIA Research and Purdue University present Scenethesis. It combines an LLM for rough scene layout with a vision module that refines object placement using image guidance, plus optimization to prevent
-

Researchers Introduce CHAI for Precise AI Video Captioning
By
–
What if AI could caption videos with the precision of a professional filmmaker? Researchers from Carnegie Mellon University and Harvard University introduce CHAI just for that. They built a structured video language using hundreds of visual primitives defined with filmmakers.
-

Researchers Introduce Laser: A More Efficient Visual Reasoning AI Model
By
–
What if visual reasoning could be 97% more efficient while outperforming top models? Researchers from MBZUAI, Fudan, Renmin, and Harvard introduce Laser. Instead of forcing step-by-step text, it uses “forest-before-trees” reasoning—maintaining a big-picture understanding
