The first disaggregated inference demo for AI agents is now live. At #COMPUTEX2026, SambaNova demonstrated premium inference running in production at VC2 — using NVIDIA B200 GPUs for prefill and SambaNova RDUs for decode. The result: 2x faster inference than B200-only
COMPUTING
-
@alphasignalai — 2026-06-03
By
–
YES, the CVE run makes that concrete, 100% accuracy at 85.1% fewer tokens, while the other systems stayed under 25%. Only word to push back on is "unprecedented" though, CodeAct was doing code-as-actions back at ICML 2024.
-
OpenAI Codex Launches Sites: Easy Web Page Creation and Deployment
By
–
Ayer OpenAI lanzó Sites como una nueva funcionalidad dentro de Codex para facilitar la creación y despliegue de páginas webs.
— Carlos Santana (@DotCSV) 3 juin 2026
Hasta aquí todo ok, pero!pic.twitter.com/fLNdqj2DQ1Yesterday OpenAI launched Sites as a new feature within Codex to facilitate the creation and deployment of web pages. So far so good, but!
-
GLM endpoint slow, mistakes help understanding, coding product fine
By
–
Yesterday evening GLM endpoint was slow, high traffic during U.S. working hours. Today it's making seemingly sillier mistakes, but every one it makes helps me understand the solution better — which has its benefits. Not perfect, but the subfrontier coding product is fine as-is.
-

BIGSET: A New Standard in AI Architecture, Beyond Simple Wrappers
By
–
But before we dive in, I have to say: BIGSET goes far beyond a simple wrapper. The tech sets a new standard for AI architecture: → AI infers the schema automatically
→ Sub-agents work in parallel
→ Source tracking per row
→ Exports to CSV/XLSX and more! keep scrolling ↓ -

The true cost of AI: every query has its price
By
–
For three years, AI was served to us as a free meal. That's over. Its real cost is not training it, but running it endlessly. Each query is a meter ticking. The token, the unit of consumption for models, becomes to AI what the barrel
-

Perplexity’s New Agent Architecture: Python, Sandbox, and Efficiency Gains
By
–


Perplexity stopped treating search as one API call. Its agents now write Python that fans out queries, filters results, and joins evidence in a sandbox. On a 200+ CVE task: 100% accuracy, 85.1% fewer tokens. The SDK is private. The pattern isn't, so lets use it with Hermes?
-
Hallo-Live: streaming audio-video avatar generation from Fudan & Baidu
By
–
What if your avatar could talk and move in real time?
— 机器之心 JIQIZHIXIN (@jiqizhixin) 3 juin 2026
Fudan & Baidu present Hallo-Live: streaming audio-video avatar generation.
Dual-stream diffusion + future-expanding attention syncs lips to upcoming speech.
Human-preference distillation preserves quality. 20 FPS, 0.94s… pic.twitter.com/0GHiOm5h2xWhat if your avatar could talk and move in real time? Fudan & Baidu present Hallo-Live: streaming audio-video avatar generation. Dual-stream diffusion + future-expanding attention syncs lips to upcoming speech. Human-preference distillation preserves quality. 20 FPS, 0.94s
-

Full-load GPU limited to 220W with DFlash, DDTree optimizations
By
–
Looks like this under full-load btw Lots of juice to squeeze yet with DFlash / DDTree / Spec. Decoding / etc Also, power limiting the GPUs to 220w down from 440w as well (okay w/ leaving the perf. loss on the table given the heat / energy savings from that)

