MoE is strictly worse for folks using discrete consumer GPUs like a 3090 AFAICT. It seems pretty good for unified memory stuff like Macs though.
TECHNOLOGY
-
Microsoft’s Mai Voice leads AI speech generation frontier
By
–
https://
copilot.microsoft.com/labs/audio-exp
ression
… Mai Voice – Leading the frontier of AI speech generation Bravo team -
Cerebras MCP Server Enables 20x Faster AI Inference
By
–
Introducing ⚡️Cerebras MCP Server ⚡️
— Cerebras (@cerebras) 28 août 2025
You can now turbocharge any AI editor that uses MCP for tool calling with 20x faster inference.
Here is Claude Code writing files at breakneck speed via Cerebras.
Get started now 👇 pic.twitter.com/Mv206wgeb2Introducing Cerebras MCP Server You can now turbocharge any AI editor that uses MCP for tool calling with 20x faster inference. Here is Claude Code writing files at breakneck speed via Cerebras. Get started now
-
OpenAI Releases Realtime Voice API
By
–
OpenAI released a production-ready Realtime API for building voice agents and a new gpt-realtime model! https://t.co/ukTLH2BdtP pic.twitter.com/pFTOeRauae
— 🚨 AI News | TestingCatalog (@testingcatalog) 28 août 2025OpenAI released a production-ready Realtime API for building voice agents and a new gpt-realtime model!
-

Grok Code Fast 1: Versatile AI Coding Model Builds Game Daily
By
–
Grok Code Fast 1 is versatile across the full stack and is particularly strong at TypeScript, Python, Java, Rust, C++, and Go.
— xAI (@xai) 28 août 2025
Using Grok Code Fast 1, @DannyLimanseta built the following game in a day. pic.twitter.com/rz2RgBno5lGrok Code Fast 1 is versatile across the full stack and is particularly strong at TypeScript, Python, Java, Rust, C++, and Go. Using Grok Code Fast 1, @DannyLimanseta built the following game in a day.
-

Grok Code Fast 1: New Lightweight Model for Speed Affordability
By
–
We built Grok Code Fast 1 from scratch, starting with a brand-new lightweight model architecture. Combined with novel improvements to accelerate serving efficiency, Grok Code Fast 1 sets a new standard for both speed and affordability.
-
xAI Model API Pricing: $0.20 per Million Input Tokens
By
–
The model is generally available via the xAI API, priced at $0.20 / 1M input tokens, $1.50 / 1M output tokens, and $0.02 / 1M cached tokens.
-
Cyberdissidents and technology-driven sabotage operations ahead
By
–
Et surtout utilisez les technologies. L'avenir est aux cyberdissidents et aux opérations de sabotage.
-

Scenario narratives guide strategic tech planning and stakeholder alignment
By
–
A scenario narrative in tech domains supports structured thinking, helps anticipate shifts, guides strategic planning, and aligns stakeholders through a concrete view of possible developments. Source @Gartner_inc Link https://
gtnr.it/4m5v9zT via @antgrasso -
Parallel Agents Emerging as Key Scaling Technique for AI
By
–
Parallel agents are emerging as an important new direction for scaling up AI. AI capabilities have scaled with more training data, training-time compute, and test-time compute. Having multiple agents run in parallel is growing as a technique to further scale and improve