
Opus 4.8 is live. Benchmarks especially significant jump in Agentic coding, but more important: „Fast mode is available for Opus 4.8. It's the same model at roughly 2.5x the speed, and we've made it three times cheaper than before.“

By
–

Opus 4.8 is live. Benchmarks especially significant jump in Agentic coding, but more important: „Fast mode is available for Opus 4.8. It's the same model at roughly 2.5x the speed, and we've made it three times cheaper than before.“
By
–
We are starting to be quite bullish about getting in the data infrastructure business.
— Julien Chaumond (@julien_c) 28 mai 2026
I just cloned 68 TB (while I only have a 4TB local disk) to my @huggingface training bucket in 1 minute 55 seconds, thanks to Xet deduplication and all our infra optimizations.
You can host… pic.twitter.com/qfm9QvaIdj
We are starting to be quite bullish about getting in the data infrastructure business. I just cloned 68 TB (while I only have a 4TB local disk) to my @huggingface training bucket in 1 minute 55 seconds, thanks to Xet deduplication and all our infra optimizations. You can host
By
–
https://t.co/b774dXzDJ7 https://t.co/NeC9DfB4rI
— Aravind Srinivas (@AravSrinivas) 28 mai 2026
Perplexity Computer can now help prepare your federal tax return. Select “Navigate my taxes” on Computer to give it a shot.
By
–
Why does this matter for AI inference specifically? Training = throughput problem. Inference = latency problem. When a user talks to an AI assistant, tokens have to return fast. Latency, memory access, bandwidth, and interconnect all matter, not just raw compute. In large AI
By
–
“This technology, this work would be big enough to require dedicated silicon… you needed to start with a clean sheet of paper and you needed to do something fundamentally different.”@andrewdfeldman joined @PeterDiamandis on MOONSHOTS to discuss wafer-scale AI compute, fast… pic.twitter.com/RNv3i2VTZL
— Cerebras (@cerebras) 28 mai 2026
“This technology, this work would be big enough to require dedicated silicon… you needed to start with a clean sheet of paper and you needed to do something fundamentally different.” @andrewdfeldman joined @PeterDiamandis on MOONSHOTS to discuss wafer-scale AI compute, fast

By
–
Want ONE BILLION free LLM tokens a month without juggling a dozen different APIs? Now you can *legally* unlock that massive inference capacity by combining the free tiers of Google, Groq, SambaNova, Mistral, and GitHub Models. The only problem is the headache of managing all

By
–
What if you could make AI language models smarter by reusing the same layers over and over? Researchers from KAIST, KRAFTON, and UC Berkeley present LoopMDM(Looped Diffusion Language Models). They selectively loop early-middle transformer layers in masked diffusion models—no

By
–
Introducing Dynamo Snapshot, our approach for fast startup for inference workloads on Kubernetes, which reduces startup time from minutes to under 5 seconds. In production inference deployments demand fluctuates over time. Cold-starting inference workloads can take minutes,
By
–
Web designers after reading this: https://t.co/yONuEtjT8L pic.twitter.com/p3y16ldruL
— Charly Wargnier (@DataChaz) 27 mai 2026
Web designers after reading this:
By
–
Codex for parallel browser-using subagents: https://t.co/Iqa3RgcBwD
— Greg Brockman (@gdb) 27 mai 2026
Codex for parallel browser-using subagents: