AI Inferencing chips. $100 billion. NVIDIA bought @GroqInc for $20 billion, which was its closest competitor. Jensen got a deal!
COMPUTING
-
Rumored Gemini Flash achieves 92% GPT-5.5 performance at 15-20x lower inference cost
By
–
Rumors about the new Gemini Flash coming in. And holy, if true then big: 92% of GPT-5.5’s coding and reasoning performance, reportedly at 15–20x lower inference cost. And the latency? Sub-200ms for most queries. That would be nuts. no joke.
-
Qwen 27B Speed Boost with TurboQuant
By
–
🚨 LOCAL AI JUST GOT A SERIOUS SPEED BOOST
— Charly Wargnier (@DataChaz) 14 mai 2026
I just found a custom setup for Qwen 27B.
> Qwen 3.6 quantized with @GoogleDeepMind's TurboQuant
> from 21 to 34 tokens per second
> +40 % performance boost 🤯
feels incredibly smooth on my Mac.
open-source code just dropped ↓ https://t.co/0zXA4uCXyzLOCAL AI JUST GOT A SERIOUS SPEED BOOST I just found a custom setup for Qwen 27B. > Qwen 3.6 quantized with @GoogleDeepMind
's TurboQuant
> from 21 to 34 tokens per second > +40 % performance boost feels incredibly smooth on my Mac. open-source code just dropped ↓ -

Google’s new Gemini model to rival GPT-5.5 at I/O
By
–
Lets go: Google’s next Gemini model is expected to compete with GPT-5.5 Google is reportedly preparing to unveil a new Gemini model at I/O, positioning it near OpenAI’s recent GPT-5.5 rather than the more elusive Anthropic Mythos. Google i/o got even more exciting
-
Microsoft Build as a Key Indicator for Enterprise AI Trends
By
–
#MicrosoftBuild continues to be one of the most important events to watch for understanding where enterprise AI, developer platforms, cloud infrastructure, and intelligent applications are heading next. The presence of @satyanadella , alongside sessions focused on GitHub
-
Optimizing Infrastructure for Agentic AI Inference
By
–
Delivering agentic inference at scale requires balancing efficiency across:
— NVIDIA AI (@NVIDIAAI) 13 mai 2026
1) Models and algorithms
2) Software
3) Compute
Our full-stack platform continuously optimizes for these inputs using extreme co-design across compute, networking, storage, and memory. Plus, software… pic.twitter.com/rzoF9wyF1NDelivering agentic inference at scale requires balancing efficiency across: 1) Models and algorithms
2) Software 3) Compute Our full-stack platform continuously optimizes for these inputs using extreme co-design across compute, networking, storage, and memory. Plus, software -
Team of agents solving a physics problem
By
–
watching a team of agents tackling a hard theoretical physics problem is quite mesmerizing – self-correcting, deriving hard equations, computing intermediate results, re-estimating the best approach https://t.co/EySjaCtUFZ pic.twitter.com/RhUmNXkGLB
— Thomas Wolf (@Thom_Wolf) 13 mai 2026Watching a team of agents tackle a challenging theoretical physics problem is quite mesmerizing—they self-correct, derive complex equations, compute intermediate results, and re-evaluate the best approach.
-
Cline SDK Revolutionizes AI Development
By
–
The new Cline SDK just dropped and it changes everything 🔥
— Charly Wargnier (@DataChaz) 13 mai 2026
→ Rebuilt agent harness
→ Fully open SDK (npm i @cline/sdk)
→ Beats Claude Code on tbench
→ Hub-backed persistent sessions
→ Cron automations
best part?
they just open-sourced the SDK so you can build on it too👀 https://t.co/3Rg0xBBnPSThe new Cline SDK just dropped and it changes everything → Rebuilt agent harness
→ Fully open SDK (npm i @cline/sdk)
→ Beats Claude Code on tbench
→ Hub-backed persistent sessions
→ Cron automations best part? they just open-sourced the SDK so you can build on it too -

Tolaria: A New Mac App for Markdown Knowledge Bases Inspired by Karpathy
By
–
Karpathy's LLM wiki idea just became a real Mac app! Tolaria is a Mac app for managing markdown knowledge bases. Your notes are plain markdown files stored in a git repository. No cloud, no subscriptions, completely offline. The core idea: knowledge bases work better when
-
The Impact of Latency on AI Deployment and Business Performance
By
–
Here’s the problem with most AI setups: → Data is sent to the cloud
→ Models process remotely
→ Decisions come back with delay That delay is latency. And in operations, latency slows performance, limits responsiveness, and impacts business value.