Cerebras Inference running Llama 70B is now so fast that it outruns GPU based inference running Llama 3B. The Wafer Scale Engine runs a model 23x larger and 8x faster for a combined 184x performance gain.
Cerebras Wafer Scale Engine runs Llama 70B 184x faster
By
–
