We need more knowledge sharing about running ML infrastructure at scale! Here's the mix of AWS instances we currently run our serverless Inference API on. For context, the Inference API is the infra service that powers the widgets on @huggingface Hub model pages + PRO users and
ML Infrastructure Scaling: AWS Instance Configuration for Inference API
By
–
