AI could understand and generate images from a single, efficient model!
— 机器之心 JIQIZHIXIN (@jiqizhixin) 8 avril 2026
Tsinghua University, Xi'an Jiaotong University, and University of Chinese Academy of Sciences present Cheers!
This unified multimodal model decouples fine image details from their core semantic meaning.… pic.twitter.com/s0MgsejA97
AI could understand and generate images from a single, efficient model! Tsinghua University, Xi'an Jiaotong University, and University of Chinese Academy of Sciences present Cheers! This unified multimodal model decouples fine image details from their core semantic meaning. This new architecture stabilizes AI's understanding while boosting image generation fidelity by selectively re-injecting those details. Cheers matches or outperforms advanced unified multimodal models in both visual understanding and generation. It notably beats Tar-1.5B on GenEval and MMBench, using only 20% of the training cost and achieving 4x token compression. Breakthrough efficiency! Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation Project: github.com/AI9Stars/Cheers Model: huggingface.co/ai9stars/Chee… Paper: arxiv.org/abs/2603.12793 Our report: mp.weixin.qq.com/s/EK6cyCJz5… 📬 #PapersAccepted by Jiqizhixin