Verified by @ArtificialAnlys, @CerebrasSystems Inference is capable of serving Llama 3.1 70B at 450 tokens/sec and Llama 3.1 8B at 1,850 tokens/sec! https://t.co/hCb9MmSvOo
— AI at Meta (@AIatMeta) 27 août 2024
Verified by @ArtificialAnlys
, @cerebras Inference is capable of serving Llama 3.1 70B at 450 tokens/sec and Llama 3.1 8B at 1,850 tokens/sec!