AI Dynamics

Global AI News Aggregator

About

SRAM+HBM Hybrid with sub-1µs latency for parallel MoE decoding

2/ Presented as the first true hybrid. Lots of on-chip SRAM (like Groq) combined with HBM (Nvidia). They claim their chip-to-chip fabric has super low latency (

→ View original post on X — @kimmonismus