I don't really use GPUs, as most of our work uses TPUs. I was just confused by @jeremyphoward 's statement that it wouldn't run on consumer GPUs, when the normal techniques of model parallelism, quantization, right compression, etc. seem like they should work just fine to run it
Running Large Models on Consumer GPUs: Parallelism and Quantization
By
–