The first release of Cerebras Inference utilized only a fraction of the wafer’s bandwidth, compute, and IO capacity. With this release, we’ve re-written and optimized everything from kernels (matmul, broadcast/reduce) to ML (speculative decoding). Model is still 16-bit and the
Cerebras Optimizes Inference Engine Across Wafer Architecture
By
–