Wait, what? @UnslothAI is starting to upload MLX Dynamic Quants! I have to test them ASAP! Thanks you really rock! 🚀 unsloth.ai/docs/models/gemma…
→ View original post on X — @huggingface, 2026-04-04 19:48 UTC

By
–
Wait, what? @UnslothAI is starting to upload MLX Dynamic Quants! I have to test them ASAP! Thanks you really rock! 🚀 unsloth.ai/docs/models/gemma…
→ View original post on X — @huggingface, 2026-04-04 19:48 UTC

By
–
PSA: models on hugging face are free to use. Always have been, always will be. If you're relying on claw, today is the best time to liberate (parts of) your agent's workflow.
→ View original post on X — @huggingface, 2026-04-04 15:41 UTC
By
–
Gemma 4 watches raw video. Understands the scene. Then prompts SAM 3 to segment and RF-DETR to track.
— Maziyar PANAHI (@MaziyarPanahi) 4 avril 2026
One AI directing two others. Fighter jets. Crowds. Aerial defense footage.
All three models running locally on a MacBook. No cloud.
What scene should I point this at next? pic.twitter.com/vNVgVloAGB
Gemma 4 watches raw video. Understands the scene. Then prompts SAM 3 to segment and RF-DETR to track. One AI directing two others. Fighter jets. Crowds. Aerial defense footage. All three models running locally on a MacBook. No cloud. What scene should I point this at next?
→ View original post on X — @huggingface, 2026-04-04 14:44 UTC
By
–
llama-server -hf ggml-org/gemma-4-26b-a4b-it-GGUF:Q4_K_M openclaw onboard –non-interactive \ –auth-choice custom-api-key \ –custom-base-url "http://127.0.0.1:8080/v1" \ –custom-model-id "ggml-org-gemma-4-26b-a4b-gguf" \ –custom-api-key "llama.cpp" \ –secret-input-mode plaintext \ –custom-compatibility openai \ –accept-risk
→ View original post on X — @huggingface, 2026-04-04 00:22 UTC
By
–
Time to move to open or local models from Hugging Face! All instructions are here: huggingface.co/blog/liberate… Boris Cherny (@bcherny) Starting tomorrow at 12pm PT, Claude subscriptions will no longer cover usage on third-party tools like OpenClaw. You can still use these tools with your Claude login via extra usage bundles (now available at a discount), or with a Claude API key. — https://nitter.net/bcherny/status/2040206440556826908#m
→ View original post on X — @huggingface, 2026-04-04 00:06 UTC
By
–
First public PR in years from the legend @JeffDean was to @huggingface transformers. Any other proud moment for us and our community! Omar Sanseviero (@osanseviero) New Jeff Dean fact: the transformers PR for Gemma 4 had 14 authors and Jeff was one of them — https://nitter.net/osanseviero/status/2040177838817476879#m
→ View original post on X — @huggingface, 2026-04-03 22:53 UTC

By
–
BREAKING: New bartowski Gemma-4 26B-A4B-it MoE GGUF Just Dropped 🤯 Dropped the IQ4_NL GGUF of Google’s Gemma-4 26B-A4B-it MoE (~26B total / ~4B active)! Bartowski’s GGUF Setup: 👇🏻 🧠Revised & quantized with llama.cpp imatrix 🚀Quant: gemma-4-26B-A4B-it-IQ4_NL.gguf 14.70 GB 💻Full 256K context window Native (text + vision) MoE Performance Wins: 👇🏻 🏆 Efficient Mixture-of-Experts (8 active / 128 total) 🚨 Clean, accurate tool calls with no overthinking 🤯 Noticeably stronger agentic workflow MoEs (beats Qwen 3.5-35B-A3B in tool-use precision) 🤖 Built for local Hermes Agent / agentic meta Quants from IQ4-XS to Q4_K_M:👇🏻 🔥IQ4_NL (14.70 GB) — best accuracy/speed balance 🔥IQ4_XS (~14.2 GB) — lightest high-quality option 🔥Q4_K_S (~15.8 GB) Q4_K_M (~17 GB) Try It 👇🏻 huggingface.co/bartowski/goo…
→ View original post on X — @huggingface, 2026-04-03 20:36 UTC

By
–
We're shipping an elaborate guide on how to profile diffusion pipelines in Diffusers to set them up for success with `torch.compile` 🔥 We devised a workflow with Claude & it turned out to be quite effective. It served its purpose well. With the help of the trace alone, we uncovered: 1. CPU <-> GPU syncs 2. CPU overheads 3. Kernel launch delays When we provided the profile trace and our observations from the trace to Claude, and helped us get rid of the issues, it did well. However, it did so iteratively. The process was intellectually fun and engaging!
→ View original post on X — @huggingface, 2026-04-03 17:07 UTC
By
–
Thanks for following us!
— Google Gemma (@googlegemma) 3 avril 2026
We're excited to see what you all build with Gemma 4!
In case you missed it, you can find all our checkpoints, with an Apache 2.0 License, on Hugging Face: pic.twitter.com/64t6fSefw4
Thanks for following us! We're excited to see what you all build with Gemma 4! In case you missed it, you can find all our checkpoints, with an Apache 2.0 License, on Hugging Face:
→ View original post on X — @huggingface, 2026-04-03 16:43 UTC

By
–
BREAKING:🚨 NVIDIA just quantized Gemma 4 31B on Hugging Face 🔥 NVFP4 compression = 4x smaller weights with frontier-level accuracy. ✅99.7% of baseline on GPQA (75.46% vs 75.71%). 📈256K context window. 🧐Multimodal (text + images + video). vLLM-ready + Blackwell optimized. VRAM requirements: ⚡️Weights only: ~16–21 GB 🚀Everyday use: Runs on 24 GB GPUs 📈Full 256K context = 32 GB VRAM sweet spot (RTX 5090-class consumer GPUs) This is the 31B-class frontier model you can actually run locally on a high-end rig. Try it today👉 huggingface.co/nvidia/Gemma-…
→ View original post on X — @huggingface, 2026-04-03 13:30 UTC