One more step towards Serverless GPU Inference Endpoints on the Hugging Face Hub now expose an option to “Automatically scale to 0” Your endpoint will be automatically paused after X minutes of not receiving requests, which will limit total costs dramatically
Hugging Face Inference Endpoints Enable Automatic Scale to Zero
By
–
