I keep seeing this advice to quantize the KVCache to 4-bit and save on memory Please don’t do that KV Cache quantization beyond FP8 usually is asking for a nerfed and incoherent model
KV Cache Quantization Beyond FP8 Degrades Model Performance
By
–
By
–
I keep seeing this advice to quantize the KVCache to 4-bit and save on memory Please don’t do that KV Cache quantization beyond FP8 usually is asking for a nerfed and incoherent model