3/ Sparse-Quantized Representation – a new compressed format and quantization technique that enables near-lossless compression of LLMs across model scales; “allows LLM inference at 4.75 bits with a 15% speedup”.
Sparse-Quantized Representation enables 4.75-bit LLM inference
By
–
