With workloads that don’t fully load the GPU you can often get a substantial throughput increase by running multiple processes on each GPU, but you need to manually enable MPS to get concurrency: https://
docs.nvidia.com/deploy/mps/ind
ex.html
…
GPU Throughput Increase with Multiple Processes Using MPS
By
–