I've just quantized CodeLlama 7b Python to 4-bit with MLX, meaning you can now run this model super fast on Apple Silicon. Here's the link to the model! https://
huggingface.co/mlx-community/
CodeLlama-7b-Python-4bit-MLX
… By the end of the day, my goal is to add all the new models. The 13B one is almost done!
CodeLlama 7B Python quantized to 4-bit on MLX
By
–