Deploying LLMs is memory inefficient and compute-intensive. At the #ACL2023 Google booth at 3pm today, @chunliang_tw will describe a new method, distilling step-by-step, that trains successful small models using fewer training data for fine-tuning and distillation.
Google Distillation Method Trains Smaller LLMs Efficiently
By
–