AI Dynamics

Global AI News Aggregator

About

Sparse-Quantized Representation enables 4.75-bit LLM inference

3/ Sparse-Quantized Representation – a new compressed format and quantization technique that enables near-lossless compression of LLMs across model scales; “allows LLM inference at 4.75 bits with a 15% speedup”.

→ View original post on X — @dair_ai