AI Dynamics

Global AI News Aggregator

About

Cerebras Wafer Scale Engine runs Llama 70B 184x faster

Cerebras Inference running Llama 70B is now so fast that it outruns GPU based inference running Llama 3B. The Wafer Scale Engine runs a model 23x larger and 8x faster for a combined 184x performance gain.

→ View original post on X — @cerebras