The model is generally available via the xAI API, priced at $0.20 / 1M input tokens, $1.50 / 1M output tokens, and $0.02 / 1M cached tokens.
LLMS
-
Grok Code Fast 1: New Fast, Economic Reasoning Model for Coding
By
–
Introducing Grok Code Fast 1, a speedy and economical reasoning model that excels at agentic coding. Now available for free on GitHub Copilot, Cursor, Cline, Kilo Code, Roo Code, opencode, and Windsurf.
-

Transforming Knowledge for LLM-First Applications and Education
By
–
Transforming human knowledge, sensors and actuators from human-first and human-legible to LLM-first and LLM-legible is a beautiful space with so much potential and so much can be done… One example I'm obsessed with recently – for every textbook pdf/epub, there is a perfect
-
MAI-1-preview: New In-House Foundation Model Released
By
–
Introducing MAI-1-preview
– our first foundation model trained end to end in house
– in public testing on LMArena
– we’re excited to be actively spinning the flywheel to deliver improved models -

Microsoft Launches MAI-Voice-1 and MAI-1-preview Models
By
–
Excited to share our first @MicrosoftAI in-house models: MAI-Voice-1 and MAI-1-preview. Details and how you can test below, with lots more to come
-
MAI-Voice-1: Expressive Natural Voice Generation Model
By
–
Introducing MAI-Voice-1
– most expressive, natural voice generation model I've ever used (might be a bit biased)
– super efficient, generating a minute of audio in <1 second on a single GPU
– live now in Copilot Daily + Podcasts
Try it in Copilot Labs too: -
Enthusiastic about series, reminder to complete LLM project
By
–
Wow, loving the series!! A reminder to self to finish build llm from scratch 😀
-
OpenAI Launches gpt-realtime Speech-to-Speech Model Updates
By
–
Introducing gpt-realtime — our best speech-to-speech model for developers, and updates to the Realtime API
-
Early Reinforcement Learning and Reasoning Chains Development
By
–
Very early days of RL, and we do see this a bit with reasoning chains.
-

Building NVIDIA Nemotron Nano Chat App with AnyCoder
By
–
vibe coding a NVIDIA-Nemotron-Nano-9B-v2 chat app in anycoder only took a couple a prompts and deployed with zero-gpu on Hugging Face NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning