Tried more quants (4-bit, 5-bit) on upstream llama.cpp Next up is benchmarking TurboQuant for the same quants The goal is finding the best quant that fits with the highest context in q8_0/turbo3 asymmetric
Benchmarking quantization levels with TurboQuant for best context fit
By
–
