Cerebras Inference perf update:
Llama3.1-8B: 1,8001,927 tokens/s
Llama3.1-70B: 450481 tokens/s
Stillfor the most popular open model in the world. https://
inference.cerebras.ai
Cerebras Inference Speed Update: Llama 3.1 Performance Benchmarks
By
–
