
Looks like DeepSeek is handing hardware optimization control directly to developers in the latest DeepGEMM update. For the fp8_mqa_logits function, the weights tensor dtype now explicitly dictates the accumulation precision.

By
–

Looks like DeepSeek is handing hardware optimization control directly to developers in the latest DeepGEMM update. For the fp8_mqa_logits function, the weights tensor dtype now explicitly dictates the accumulation precision.

By
–
14x RTX 3090s + Qwen 3.6 27B Running 42 agents IN PARALLEL at full 256k context – exl3 6bpw
– fp8 KV Cache
– Aphrodite Inference Engine w/ tp=2, pp=7 The world of agents will run locally btw
By
–
https://t.co/LizAzIh7dN pic.twitter.com/VVreg7R7OW
— Bojan Tunguz (@tunguz) 2 juin 2026
At Microsoft Build today, Mustafa Suleyman predicted three more OOM jumps in the amount of training compute between now and summer 2029.

By
–
From unboxing to AI agent in minutes. Getting an agent running used to mean sourcing a model, configuring an inference backend, installing a runtime, and wiring it all together. The new NemoClaw install path on DGX Spark replaces that with a single command. DGX Spark also
By
–
Yep! Codex on my Mac Mini, mostly controlled via Codex Mobile
By
–
The AI era needs a new CPU.
— NVIDIA (@nvidia) 2 juin 2026
Meet NVIDIA Vera, 80% faster agentic task completion than x86.
Built for AI factories. Built for what's next. The CPU for agents has arrived.
🔗 https://t.co/Qe1S4f4Qz5 pic.twitter.com/p7nh6ot6r8
The AI era needs a new CPU. Meet NVIDIA Vera, 80% faster agentic task completion than x86. Built for AI factories. Built for what's next. The CPU for agents has arrived. https://
nvda.ws/4ugo9Uf
By
–
Today we're announcing that hybrid agentic inference is coming to Perplexity Computer.
— Perplexity (@perplexity_ai) 2 juin 2026
Computer can split tasks between a local model running on your machine and frontier models in the cloud. This keeps private data on your device and maximizes token efficiency.
Coming soon. pic.twitter.com/6t3PrmI1FX
Today we're announcing that hybrid agentic inference is coming to Perplexity Computer. Computer can split tasks between a local model running on your machine and frontier models in the cloud. This keeps private data on your device and maximizes token efficiency. Coming soon.

By
–



This came as a surprise: Microsoft has unveiled handheld and desktop devices designed to control one's agents. It reminds me of what I had expected from OpenAI’s hardware-standalone devices for controlling agents.

By
–
Badge form factor AI device With camara, voice control and AI Agents

By
–
Project Solara A new chip for an Agent-first world and Agent-first devices.