Just do this: brew install llama.cpp –HEAD Then; llama-server -hf ggml-org/gemma-4-26B-A4B-it-GGUF:Q4_K_M
SOFTWARE
-

Google Releases Gemma 4 in Four Variants
By
–

BREAKING : Google released Gemma 4 in 4 different variants: 31B, 26B MoE, 2B and 4B! Offline on-device AI apps got a huge upgrade!
-

Music Creation and Sharing App Built with Replit Canvas
By
–
An app to make and share your music anywhere Made with Replit Canvas
-
Perplexity Computer Now Helps Prepare Federal Tax Returns
By
–
Perplexity Computer can now help prepare your federal tax return.
— Perplexity (@perplexity_ai) 2 avril 2026
Select “Navigate my taxes” on Computer to give it a shot. pic.twitter.com/XppQTXz4JWPerplexity Computer can now help prepare your federal tax return. Select “Navigate my taxes” on Computer to give it a shot.
→ View original post on X — @aravsrinivas, 2026-04-02 16:25 UTC
-

Gemma4 vindicated: dense models triumph over mixture of experts
By
–
Gemma4 is amazing. You'll read that everywhere. Let's focus on what is HUGE here: the revenge of dense models…. Throw away your b200, not needed anymore, throw away the millions of lines of code we had to write to make MOEs faster, training stable etc… throw away your router-aware kernel, your EP DEEP GEMM, throw away the auxiliary loss function. Welcome to simplicity, dense is the new king. FINALLY hating MoEs is back to being chad. For those who know me: I was always a moe doomer
→ View original post on X — @jeremyphoward, 2026-04-02 16:23 UTC
-
Google Gemma 4 Now Available on Modular Cloud Platform
By
–
Google Deep Mind's impressive fully-open Gemma 4 is live day-zero on Modular Cloud. Modular provides the fastest performance on NVIDIA Blackwell and AMD MI355X, thanks to MAX and Mojo🔥. The team took this impressive new model to production inference in days.🚀
→ View original post on X — @jeremyphoward, 2026-04-02 16:15 UTC
-

Gemma 4 Launch: New Open Models Across Multiple Sizes
By
–
Excited to launch Gemma 4: the best open models in the world for their respective sizes. Available in 4 sizes that can be fine-tuned for your specific task: 31B dense for great raw performance, 26B MoE for low latency, and effective 2B & 4B for edge device use – happy building!
→ View original post on X — @demishassabis, 2026-04-02 16:08 UTC
-
Understanding Open Models: Gemma’s Public AI Systems Explained
By
–
And just in case you’re wondering, "..What’s an open model?", we’ve got you covered: Basically, open models are AI systems where the model weights are publicly available for anyone to download, study, fine-tune and use on your own hardware (phones, computers, etc.). Open models can live on your hardware where your data is completely private and never has to leave your machine. Once you download an open model onto your device, it can run anywhere regardless of internet connection or access to data centers. To name a few examples, Gemma models can run in your pocket, underwater, in outer space, from subway tunnels, and on high-altitude flights without needing a cell tower or WiFi signal. As base models are released (the 'blueprints'), people can then further modify them for specialized use cases via fine-tuning. We’ve seen this in the Gemmaverse, where developers have downloaded Gemma over 400 million times and built more than 100,000 variants. Have you used an open model before? Let us know if you have any other questions about this neat technology!
-
Building Systems Through Conversation with AI Intelligence
By
–
It is magic that I can build a system/site like this just by talking with a digital intelligence running on melted sand.
-
Mistral TTS Local Deployment Challenges Explained
By
–
Yeah dude, that latest Mistral TTS model is a major pain in the a** to figure out how to run locally.
