We just shipped Tier 3 limits in the Gemini API (2x-6x higher limits), available now for self serve upgrade if you have spent >=$1K with Google Cloud. You can upgrade in AI Studio (on the API key page) to Tier 3 and start scaling with Gemini right now!
@officiallogank
-
Google Gemini API Rate Limits and 429 Error Handling
By
–
yes, see here: https://
ai.google.dev/gemini-api/doc
s/rate-limits
… and the 429 error says exactly what type of limit you are hitting (we shipped this a couple weeks ago) -
Caching reduces inference costs by 4x efficiency gains
By
–
yes, caching is 4x less expensive than regular inference
-
Cache Size Impact on System Latency Performance
By
–
yes, the larger the cache, generally more latency you will see reduced, for small caches it will make less of a difference
-
Caching API Explores Alternative Implementation Strategies
By
–
Our caching API is explicit caching right now, but we are looking into other ways of doing it!
-
Gemini API Caching Documentation and Pricing Guide
By
–
Read more in the Gemini API docs: https://
ai.google.dev/gemini-api/doc
s/caching
… And check out the caching prices: https://
ai.google.dev/gemini-api/doc
s/pricing
… -
Gemini API context caching updates support 2.5 Pro Flash
By
–
Context caching updates in the Gemini API: – Added support for 2.0 Flash
– Added support for 2.5 Pro Preview
– Reduced min context size from 32K down to 4K Much more to come still, please send any feedback on the experience! -
API Returns Thought Token Count Instead of Thinking Process
By
–
The API doesn’t return thoughts, but it does tell you how much thinking it did (thought tokens).
-
Product Reaches General Availability for Production Use
By
–
Basically GA now, you can use in prod, nothing really changing except more new features
-
Smart People and Compute Resources Will Ensure Success
By
–
They will be fine, lots of smart people and compute there
