Those 20 times are 22 GPU-hours against 1.14. Throughput on one 5090 was about 8.4K tokens/s against 124K.
Full Breakdown:
GPU-hours and token throughput comparison for AI workloads
By
–
By
–
Those 20 times are 22 GPU-hours against 1.14. Throughput on one 5090 was about 8.4K tokens/s against 124K.
Full Breakdown: