lots of wafers = lots of mem for kv cache
multiple users at a time overlapping the pipeline
Wafer Memory Optimization for LLM Inference Parallelization
By
–
By
–
lots of wafers = lots of mem for kv cache
multiple users at a time overlapping the pipeline