Okay that's good to know 🙂
I was just following the GPT-3 paper numbers in the table, but like I mentioned it's possible the settings are way too conservative.
Few more things to try: we want to increase the batch size as much as possible to get higher tok/s. Are you using -r 1
Optimizing batch size and throughput for GPT-3 training
By
–