Wanna replace Anthropic/OpenAI? START WITH THIS The bible for running LLMs locally is now available online to read for free Covers what to use on – Laptop / edge / odd hardware
– Mac-first workflows
– Single RTX GPUs
– 2-4+ NVIDIA / CUDA GPUs
– General production serving
–
HARDWARE
-

Free online guide for running LLMs locally on all hardware
By
–
-
Skepticism about OSS owning high-end model hardware
By
–
Those of you who are putting all your hope in OSS – you really think that if they are serious about all of this they’ll let you own the hardware that can run the top end models?
-

MusaCoder: AI framework for generating efficient GPU kernels
By
–
Can AI write efficient GPU code from scratch? Moore Threads AI presents MusaCoder, a full-stack training framework for native GPU kernel generation. It combines progressive data synthesis, rejection fine-tuning, and execution-feedback reinforcement learning. To stabilize
-

Microsoft’s Singularity: 100,000 GPU Planet-Scale AI Infrastructure
By
–



Singularity: 100,000 GPUs. AI Infrastructure at Planet-Scale! Singularity is Microsoft’s distributed scheduling AI infrastructure, designed for the highly efficient and reliable execution of deep learning training and inference. Maximizing accelerator utilization without
-
Cerebras to share more Gemma 4 31B information updates
By
–
We will update more Gemma 4 31B information as we go! https://
inference-docs.cerebras.ai/models/gemma-4
-31b
… -
The defining metric of the 21st century: Intelligence per watt
By
–
The defining metric of the 21st century is Intelligence per watt. (IPW)
— Nina Schick (@NinaDSchick) 26 juin 2026
The throughput of Intelligence divided by the power consumed to produce it. pic.twitter.com/wcixNgdA3UThe defining metric of the 21st century is Intelligence per watt. (IPW) The throughput of Intelligence divided by the power consumed to produce it.
-
Offline voice assistant on tiny Axelera AI Mini PC
By
–
A full voice assistant, running with no internet connection at all, on a tiny, self-contained device you can put anywhere. This is the new Axelera AI Mini PC running Llama 3.2 1B as the language model, with separate speech-to-text and text-to-speech models alongside it, all on… pic.twitter.com/MUXkSdlJ7P
— Axelera AI (@AxeleraAI) 26 juin 2026A full voice assistant, running with no internet connection at all, on a tiny, self-contained device you can put anywhere. This is the new Axelera AI Mini PC running Llama 3.2 1B as the language model, with separate speech-to-text and text-to-speech models alongside it, all on
-

Local open-weight LLMs: 30B MoE sweet spot at 40 tok/sec
By
–
Have been taking different local open-weight LLMs for a test drive in different harnesses (Qwen-Code, Codex, Claude Code). 30B Mixture-of-Expert models are kind of a nice sweet spot and can solve challenging problems. And they get roughly 40 tok/sec on a Mac or DGX Spark, which
-
Heterogeneous disaggregated inference is the future of AI
By
–
Training builds models. Inference builds businesses. At @deeptechweek SF, our Chief Product & Strategy Officer Abhi Ingle shared why heterogeneous, disaggregated inference is the future of AI, and why "more intelligence per joule" is the metric that matters. @LipBuTan1
-
Hunt an RTX 3090 to run Qwen 3.5 27B now
By
–
I am not kidding, now is the time more than ever to hunt an RTX 3090 and learn how to run Qwen 3.5 27B