According to the model card gpt-oss-120b took 2.1 million H100-hours and the 20b model 1/10th of that I found quotes for H100 pricing ranging from $2/hour to $11/hour, which would indicate that the 120b model cost between $4.2m and $23.1m and the 20b between $420,000 and $2.3m
AI HARDWARE
-
Groq’s Real Value: Token Throughput Over Inference Speed
By
–
IMO this is the big opportunity for Groq, and one of the main reasons I invested. Faster inference is cool, but the real value is going to be in doing more tokens per prompt.
-
Groq AI Enables Browser and Code Interpreter Capabilities
By
–
Yes they can — here's it using a browser and code interpreter in @GroqInc! pic.twitter.com/2i46HCBLHw
— Matt Shumer (@mattshumer_) 5 août 2025Yes they can — here's it using a browser and code interpreter in @GroqInc
! -

Remarkable Performance of Compact Open-Source AI Models
By
–
What’s absolutely amazing to me is the size of these models:
• gpt-oss-120b runs on a high-end MacBook or on a single H100.
• gpt-oss-20b runs on consumer hardware. And yet, their evals are remarkable! Matching or exceeding o4-mini and o3-mini respectively in core benchmarks. -
Two AI Models Built for Agentic Tool Use and Structured Outputs
By
–
We learned how much features like agentic tool use and Structured Outputs matter to you. But also that one size wouldn’t fit all. That’s why we built two models, each designed for different needs and hardware, and applied our most rigorous safety testing.
-

Model now available on Cerebras for supersonic execution speeds
By
–
Buah, el modelo ya está disponible en Cerebras, lo que significa que lo podéis ejecutar a velocidades supersónicas!
-

Server Infrastructure Challenges for Large Model Weights Distribution
By
–
Please don’t download the weights all at once or our servers will melt
-

Open-Source LLM Matches OpenAI o4-mini on Benchmarks
By
–
gpt-oss-120b matches OpenAI o4-mini on core benchmarks and exceeds it in narrow domains like competitive math or health-related questions, all while fitting on a single 80GB GPU (or high-end laptop). gpt-oss-20b fits on devices as small as 16GB, while matching or exceeding
-

Large AI Models Executable with 80GB and 16GB VRAM
By
–
El modelo grande es ejecutable con 80GB de VRAM (poco asequible para la mayoría) y es el modelo comparable con o4-mini. El modelo mediano es ejecutable con 16GB de VRAM (bastante apto para muchos PCs) y es comparable a o3-mini. La verdad, que estén regalando modelos tan
-
Open-Source o3 Model Reaches 500 Tokens Per Second on Groq
By
–
It's over. OpenAI just crushed it.
— Matt Shumer (@mattshumer_) 5 août 2025
We have their o3-level open-source model running on @GroqInc at 500 tokens per second.
Watch it build an entire SaaS app in just a few seconds.
This is the new standard. Why the hell would you use anything else?? pic.twitter.com/ROAFtF7s3lIt's over. OpenAI just crushed it. We have their o3-level open-source model running on @GroqInc at 500 tokens per second. Watch it build an entire SaaS app in just a few seconds. This is the new standard. Why the hell would you use anything else??