Enjoyed this experiment — 128K context is great but you don’t get proportionally more attention. Each output token can only consider so much. Translation — fine
Extraction, doc Q&A — maybe?
Parity bit of input — absurd; the model is not a god
128K context: no proportional attention increase
By
–
