AI Dynamics

Global AI News Aggregator

About

DiffusionGemma generates 256 tokens in parallel, exceeds 1000 tok/s

How DiffusionGemma works: → Generates 256 tokens in parallel during each forward pass
→ Repeatedly refines the entire response
→ Exceeds 1,000 tokens per second on a single NVIDIA H100 Speed is the obvious headline. It is not the interesting part.

→ View original post on X — @ronald_vanloon