I have a good first impression, but it's still a tad slow for me (using llama.cpp). About 2x slower than gpt-oss 120B on the same hardware. I think I need to look for the NVIDIA-optimized stack.
Impressions initiales positives, mais performance ralentie avec llama.cpp
By
–