I keep seeing the same thing People with multi-GPUs wondering why local LLMs have slow performance …Because you're using the wrong Inference Engine and it's processing things one GPU at a time This old writeup of mine covers Inference Engines & Tensor Parallelism, go read it
@theahmadosman
-

DeepSeek V4 Flash runs on 4x DGX Spark cluster with Codex Cli
By
–
DeepSeek V4 Flash is now running on 4x DGX Spark / GB10 cluster Had to patch several things in vLLM to get it up w/ PyTorch fallbacks Targeted kernel optimization is next up P.S. Codex Cli w/ GPT-5.5 XHIGH handled the whole thing on its own, now we optimize those GB10 kernels x.com/TheAhmadOsman/…
-

Qwen 3.6 27B runs on single RTX 5090, crowned at-home flagship
By
–
This prediction was right, and on a single 5090 not even RTX PRO 6000 Qwen 3.6 27B is the at-home flagship model right now
-

The Underestimated Power of Running AI Models Locally
By
–
People seriously underestimate the importance of being able to run their models locally
-

GPT 5.5 and new models added to Codex CLI
By
–
GPT 5.5 (and a bunch of random model names) have been added to Codex Cli GPT 5.5 has arrived
-

Kimi K2.6 near top models, frontier AI runs locally.
By
–

Kimi K2.6 beats Opus 4.6 and is only 3 points behind all top 3 models Blows my mind that I am running this model locally on my own hardware Frontier intelligence at home is already a reality, opensource AI will win
-

Kimi K2.6 is the new open-source SOTA, says local user
By
–

Been running Kimi K2.6 locally all morning I can say with full confidence that this is the current new opensource SOTA Amazing work by Moonshot folks
-

LLM Inference Engine Stack Bottlenecks Cheatsheet
By
–

LLM Inference Engine Stack Breakdown and Workload/Bottlenecks Cheatsheet From the upcoming Inference Engine Comprehensive Article I am writing
-
Alpha in using Codex CLI to teach yourself anything
By
–
There’s so much alpha in opening Codex Cli and telling it to teach you how to do X X can be – shell basics
– optimizing local ai inference
– hacking your pet feeder camera – setting up a self-hosted search engine – creating Skills out of your workflows and past conversations
