Really tiny. We released GGUF quants, so with 8-bit precision and 4k context the 350M will use ~380 MB and the 1.2B ~1.4 GB (to give you a rough estimate).
By
–
Really tiny. We released GGUF quants, so with 8-bit precision and 4k context the 350M will use ~380 MB and the 1.2B ~1.4 GB (to give you a rough estimate).