Now available in addition to GPT 5.5 and Claude. Check out http://
alphaXiv.org!
COMPUTING
-
AlphaXiv.org now offers GPT 5.5 and Claude models
By
–
-
Perplexity Computer announces hybrid agentic inference with local and cloud
By
–
Today we're announcing that hybrid agentic inference is coming to Perplexity Computer.
— Perplexity (@perplexity_ai) 2 juin 2026
Computer can split tasks between a local model running on your machine and frontier models in the cloud. This keeps private data on your device and maximizes token efficiency.
Coming soon. pic.twitter.com/6t3PrmI1FXToday we're announcing that hybrid agentic inference is coming to Perplexity Computer. Computer can split tasks between a local model running on your machine and frontier models in the cloud. This keeps private data on your device and maximizes token efficiency. Coming soon.
-

Microsoft unveils agent control devices, reminiscent of OpenAI’s expected hardware
By
–



This came as a surprise: Microsoft has unveiled handheld and desktop devices designed to control one's agents. It reminds me of what I had expected from OpenAI’s hardware-standalone devices for controlling agents.
-

SambaNova unveils disaggregated inference cloud with Intel Xeon at COMPUTEX2026
By
–
The future of inference is disaggregated. At #COMPUTEX2026, Vector Core Compute unveiled the world's first fully disaggregated inference cloud powered by Intel Xeon and SambaNova RDUs—designed for the scale, speed, and economics modern AI demands. Excited to help bring this
-
AI costs soar like AC bills, contradicting Altman’s cheap forecast
By
–
Last June, Sam Altman wrote a blog post predicting “Intelligence too cheap to meter is well within grasp” and the “cost of intelligence should eventually converge to near the cost of electricity.” A year later, AI costs feel more like an AC bill in the middle of a heat wave.
-
RTX Spark running 120B parameter model locally
By
–
RTX spark running 120b parameter model locally. Ngl, pretty cool
-

First disaggregated inference cloud VectorCore Compute launched at COMPUTEX2026
By
–
At @LipBuTan1
's #COMPUTEX2026 keynote today, @RodrigoLiang stepped onstage with @RFS_Vista to power up the world's first disaggregated inference cloud, VectorCore Compute (VC2), launched by @Vista_Equity and Cambium Capital. Three chips ran disaggregated inference, live from the -
ElevenLabs previews on-device Text-to-Speech at Warsaw Summit
By
–
At the ElevenLabs Summit in Warsaw, we previewed on-device Text to Speech – a new model architecture that delivers human-level quality on limited hardware without an internet connection. pic.twitter.com/iZuztsIR9N
— ElevenLabs (@ElevenLabs) 2 juin 2026At the ElevenLabs Summit in Warsaw, we previewed on-device Text to Speech – a new model architecture that delivers human-level quality on limited hardware without an internet connection.
-

GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization
By
–
GPU Forecasters Language Models as Selective Surrogates for Kernel Runtime Optimization
-

Representation Forcing for Bottleneck-Free Unified Multimodal Models
By
–
Most Unified Multimodal Models still generate images through a frozen VAE, which means perception and generation are not fully learned in one model. This paper fixes this by making the decoder first predict