ALiBi promises fast, long extrapolation but in production models only extrapolate 10-20% longer than its training context. This is why ALiBi models such as BTLM-3B-8K and MPT-7B-8K are trained on a mix of 2K and 8K contexts – ALiBi on its own cannot extrapolate from 2K to 8K.
ALiBi Context Extrapolation Limits in Production Language Models
By
–