I keep seeing the same thing People with multi-GPUs wondering why local LLMs have slow performance …Because you're using the wrong Inference Engine and it's processing things one GPU at a time This old writeup of mine covers Inference Engines & Tensor Parallelism, go read it
HARDWARE
-
Microsoft’s Sense Cam: Always-Worn Camera Technology Evolution
By
–
Back in 1995 Microsoft had a research group in San Francisco that was wearing cameras all day long, discovering how an always-worn camera would change human life. I interviewed that team long ago. They called their camera “Sense Cam.”
— Robert Scoble (@Scobleizer) 25 avril 2026
Today @Looki_ai came to my house to give me… pic.twitter.com/PL21yXOuO6Back in 1995 Microsoft had a research group in San Francisco that was wearing cameras all day long, discovering how an always-worn camera would change human life. I interviewed that team long ago. They called their camera “Sense Cam.” Today @Looki_ai came to my house to give me
-

Telepathic Control of Robot Arms: Brain-Machine Interface Innovation
By
–
Telepathic control of robot arms https://t.co/oX21dgBA1z
— Elon Musk (@elonmusk) 24 avril 2026Telepathic control of robot arms
-

DeepSeek-V4-Pro Performance Benchmarks on NVIDIA Blackwell
By
–
Day 0 performance is here: DeepSeek-V4-Pro running on NVIDIA Blackwell Ultra. Using @vllm_project
's Day 0 recipe, we’ve captured the initial performance Pareto for DeepSeek’s flagship 1M long-context model. This curve highlights the baseline for balancing AI factory -

Starship becomes most powerful moving object ever constructed
By
–
Starship is the most powerful moving object ever made pic.twitter.com/8FpSmjtPo2
— Elon Musk (@elonmusk) 24 avril 2026Starship is the most powerful moving object ever made
-
Google Cloud Next Unveils TPUs, Gemini Enterprise Agent Platform
By
–
Our teams have been busyyy! Here are some key updates from the past week: — @GoogleCloud unveiled a suite of AI innovations at our Cloud Next event, including our eighth generation TPUs (TPUt for inference + TPUi for reasoning), Gemini Enterprise Agent Platform, Agentic Data
-

DeepSeek V4 Flash runs on 4x DGX Spark cluster with Codex Cli
By
–
DeepSeek V4 Flash is now running on 4x DGX Spark / GB10 cluster Had to patch several things in vLLM to get it up w/ PyTorch fallbacks Targeted kernel optimization is next up P.S. Codex Cli w/ GPT-5.5 XHIGH handled the whole thing on its own, now we optimize those GB10 kernels x.com/TheAhmadOsman/…
-
DeepSeek-V4-Pro 1.6T Model Now Available on NVIDIA Blackwell
By
–
Happy Friday!
— NVIDIA AI (@NVIDIAAI) 24 avril 2026
We just put DeepSeek-V4-Pro up on https://t.co/es07MrTxSs. It’s the world’s largest open source model at 1.6T parameters, and you can run it for free running on NVIDIA Blackwell GPUs.
Try the NVIDIA NIM API → https://t.co/zeWX4Y7Ipd pic.twitter.com/lNFsziIts4Happy Friday! We just put DeepSeek-V4-Pro up on http://
build.nvidia.com. It’s the world’s largest open source model at 1.6T parameters, and you can run it for free running on NVIDIA Blackwell GPUs. Try the NVIDIA NIM API → https://
build.nvidia.com/deepseek-ai/de
epseek-v4-pro?ncid=so-twit-300913
… -

Developer Flying to Collaborate on Next Generation Local AI
By
–
lol these headlines Btw the meta thing is that he's flying to work with the llamacpp team to unlock the next generation of local AI so he might fly back with an even cooler setup!

