The paper explores long context performance in detail. We found BTLM-3B-8K outperforms other 7B-8K models despite being trained with less than a fifth the pretraining compute and less than half the size.
BTLM-3B Outperforms Larger 7B Models in Long Context
By
–
