AI Dynamics

Global AI News Aggregator

About

GGUF Quantization: 350M Model Uses Just 380MB

Really tiny. We released GGUF quants, so with 8-bit precision and 4k context the 350M will use ~380 MB and the 1.2B ~1.4 GB (to give you a rough estimate).

→ View original post on X — @maximelabonne