AI Dynamics

Global AI News Aggregator

About

Token Generation Speed Comparison: OSS and Nemotron Models Performance

Yeah. Not quite as bad for me, but:
gpt-oss-120b: 42.22 tokens/sec
nemotron-3-super: 20.43 tokens/sec on the DGX Spark. But this is ollama and might be an implementation issue. I have yet to try the Nvidia-optimized llama.cpp version (they had one for Nemotron Nano back then)

→ View original post on X — @rasbt