On TPUs you can't use dynamic sizes in a loop, so you use mask and static sizes, and therefore the first step of the loop is as costly as the last one. On GPUs you could have a faster first step, but the cost of attention reduction is relatively low for large models.
@arthurmensch
-
Fixed Size Cache Initialization in Sampling Operations
By
–
Because you typically init caches with fixed size before sampling
-
DeepMind LLM Research at NeurIPS Discussion and Recruitment
By
–
At NeurIPS this week, reach out if you want to discuss our work around LLM at @DeepMind (Chinchilla, Retro, Flamingo), and if you're interested in working with us!