TIL you can run GPT-OSS 20B on a phone! This is on Snapdragon phones with 16GB or more of GPU-accessible memory – I didn't realize they had the same unified CPU-GPU memory trick that Apple Silicon has (The largest iPhone 17 still maxes out at 12GB, so not enough RAM to run https://
x.com/nexa_ai/status
/nexa_ai/status/1975232300985291008
…
COMPUTING
-
Running GPT-OSS 20B on Snapdragon Phones with GPU Memory
By
–
-
Snapdragon Phones Now Support 16GB GPU-Accessible RAM
By
–
Whoa! 16GB of RAM needed – I hadn't realized there were Snapdragon phones out there with that much GPU-accessible RAM, so TIL! https://
x.com/nexa_ai/status
/nexa_ai/status/1975233621691924962
… -
Building Declarative Data Pipelines with AI Analytics on Databricks
By
–
[Demo] Learn how to build a Lakeflow Declarative Pipeline on Databricks Free Edition using real-time flight data from OpenSky!
— Databricks (@databricks) 10 octobre 2025
See how to ingest billions of avionics events, build declarative pipelines, and apply AI-driven analytics with natural language queries for instant… pic.twitter.com/c0FmLWUMiq[Demo] Learn how to build a Lakeflow Declarative Pipeline on Databricks Free Edition using real-time flight data from OpenSky! See how to ingest billions of avionics events, build declarative pipelines, and apply AI-driven analytics with natural language queries for instant
-

Data Centre NIMBYism: Energy and Political Backlash Emerging
By
–
I think @nathanbenaich is spot on here about Data Centre NIMBYism hitting the political mainstream. Just wait until the ‘environmental’ lobby and local communities clock on to the scale of compute infrastructure and what that means for energy, energy prices and the scale of
-

Performance Optimization and Cost Reduction for Production AI Systems
By
–
And the most compelling thing is getting this performance while ALSO reducing costs when it’s time to take the system to production and things need to scale.
-

GPT-5 Outperforms Mini and Nano in Multi-Agent Tool Tasks
By
–
Our team tested single- vs multi-agent setups using GPT-5, GPT-5-mini, and GPT-5-nano, using 10K+ tools across 30 domains. In the single-agent setup, GPT-5-mini and GPT-5-nano performance degrades with longer context and more reasoning, while GPT-5 remains remarkably consistent.
-

Processing 1.3 Quadrillion Tokens Monthly Demonstrates Massive Scale
By
–
We processed over 1.3 Quadrillion tokens last month – that's 1,300,000,000,000,000 tokens! or to put it another way that's 500M tokens a second or 1.8 Trillion tokens an hour…
-

Meta’s Stochastic Activation Achieves 90% Sparsity LLM Speedup
By
–
First ever activation function swapping strategy just dropped for LLMs! Meta’s new Stochastic Activation proposes a new method which randomly selects between non-linear functions (ReLU & SiLU) in the FFN of LLM, achieving an activation sparsity of 90% with 1.65x CPU speedups!
-
NVIDIA InferenceMAX Delivers $75M Revenue Potential Per System
By
–
To help companies get the most value, NVIDIA systems are built to deliver as much output as possible at AI factory scale. Recent InferenceMAX v1 results show NVIDIA sets the standard:
— NVIDIA (@nvidia) 10 octobre 2025
One NVIDIA system can enable $75 million in revenue for AI companies using (tag DeepSeek AI)… https://t.co/mWYKsySFXATo help companies get the most value, NVIDIA systems are built to deliver as much output as possible at AI factory scale. Recent InferenceMAX v1 results show NVIDIA sets the standard: One NVIDIA system can enable $75 million in revenue for AI companies using (tag DeepSeek AI)
-

Multimodal Edge AI: Vision Audio Integration for Speaker Identification
By
–
Combining vision + audio on edge AI to identify who's speaking. A huge number of great applications across all kinds of industries and use cases. Real world engineering at its best Demo: https://
eu1.hubs.ly/H0nKK6B0