10/ Quip – compresses trained model weights into a lower precision format; combines lattice codebooks with incoherence processing to create 2 bit quantized models; significantly closes the gap between 2 bit quantized LLMs and unquantized 16 bit models.
Quip: 2-Bit Quantization for Large Language Models
By
–
