2M token context sounds incredible but I wonder how it works in practice. KV cache at that scale is a real engineering problem, and results are quite often disappointing for higher context, especially for inter-connected questions that basically needs some sort of "retrieval"
2M Token Context: KV Cache Engineering Challenges and Limitations
By
–