Can LLMs run on ultra-low-bit memory without tanking accuracy? Researchers from Together AI, University of Sydney, and UIUC present OSCAR — a method that uses offline, attention-aware covariance analysis to design fixed rotations and clipping thresholds for 2-bit KV cache
OSCAR: 2-bit KV cache for LLMs without accuracy loss
By
–
