Focused on making long context better then extending to 2M again
LLMS
-
vLLM Prefix Caching: Optimization Technique Explained
By
–
here's my favorite, the explanation of vLLM prefix caching: http://
docs.vllm.ai/en/v0.8.5/desi
gn/automatic_prefix_caching.html
… -
Claude 2.5 Pro thinking feature minimum token limit details
By
–
Yes, the only model you can’t turn thinking off on rn is 2.5 pro, which goes down to 128 tokens, will be zero in the future revs if the model, but for now that’s the min.
-
Gemini 2.5 Flash Lite Model Speed and Capabilities
By
–
Check out the speed and capabilities of our new Gemini 2.5 Flash Lite model ⬇️ https://t.co/vmOPUzyJ7t
— Jeff Dean (@JeffDean) 17 juin 2025Check out the speed and capabilities of our new Gemini 2.5 Flash Lite model
-

Gemini 2.5 Pro and Flash Models Now Generally Available
By
–
Very exciting to see our Gemini 2.5 Pro and 2.5 Flash models become "generally available" (which gives long term support commitments without changing the model), as well as a preview release of a new 2.5 Flash Lite model that offers very low latency & pricing for many uses!
-
Mixing English and Python for Seamless LLM Interaction
By
–
We’ve been working towards a similar goal in a different way. We still use English to talk to the LLM and python for code, but we’re working on a system where you can mix the two much more conveniently.
-

O3 AI Model Responses to September 2024 Problems
By
–
Estos ejemplos que he compartido son las respuestas de o3/o3 pro a problemas de septiembre 2024. Las respuestas a los nuevos prompts que me estáis compartiendo las tenéis como respuestas del primer tweet. Se agradece la difusión con un RT!
-
Testing o3 Pro with Complex Problem Formulation
By
–
Dame material para meterle a o3 pro y probamos todo esto! Que si depende de mi redactar un problema con lo que me estás diciendo…
-

Google Announces Gemini 2.5 Series Model Release
By
–
It’s been an amazing few months of relentless building, shipping, and optimising our models incorporating your feedback. Excited for more users and developers to try out the incredible Gemini 2.5 series!