AI Dynamics

Global AI News Aggregator

About

Burning model into chip achieves 51k tokens per second

now add this to silicon that burns the model into the chip. And we will go from 17.000 token/s to 51.000 tokens/s inference throughput will go on to expand so much faster than we ever could have predicted. This will make for the most absurd applications.

→ View original post on X — @linusekenstam