Wow, its happening! huggingface.co/unsloth/model… @UnslothAI can we get Gemma 26A4 and Qwen 35A3 MoE in 2-bit? (or smaller) 🙏 Is it possible to optimize Streaming Experts for size instead of trying to fit the whole model into, say, a 2-bit budget for Flash-MoE use case?
→ View original post on X — @clementdelangue, 2026-04-04 21:48 UTC