Checking with team to see if we can share or if it’s dynamic. On explicit API, you set the TTL yourself, hang tight.
@officiallogank
-
Making AI Models More Token Efficient
By
–
We are working to make them much more token efficient, stay tuned 🙂
-
New pricing model for reasoning-enabled AI language models
By
–
As mentioned at the bottom, we chose to have a “thinking off” option with a much lower price specifically so that devs moving from 2.0 had a more clear migration path. But with reasoning is definitely more expensive, very different class of models.
-
Flash 2.0 Model Reaches Stable GA Release Status
By
–
what changed on 2.0 flash? the model is stable and GA
-
Automatic Rollout: No Developer Setup Required
By
–
Automatic! Should be fully rolled out, no dev setup required
-
Explicit Caching API Reduces Costs for AI Applications
By
–
If you want to guarantee cost savings, you can continue to use the explicit caching API we shipped last May. Also, make sure to keep the initial content of the requests the same if you want them to hit the cache. More details on the launch here:
-
Google Ships Implicit Caching for Gemini API Cost Savings
By
–
We just shipped implicit caching in the Gemini API, automatically enabling a 75% cost savings with the Gemini 2.5 models when your request hits a cache We also lowered the min token required to hit caches to 1K on 2.5 Flash and 2K on 2.5 Pro!