AI Dynamics

Global AI News Aggregator

About

MiniMax Teases Sparse Attention Architecture for M3 with Significant Speedup Benchmarks

MiniMax just teased their Sparse Attention architecture for M3. The benchmarks show 9.7x prefilling speedup and 15.6x decoding speedup at 1M tokens vs M2. MiniMax deliberately went back to full attention for M2 because efficient attention wasn't production-ready. Their pretrain

→ View original post on X — @kimmonismus