2/ Effective Long-Context Scaling with LLMs – propose a 70B variant that can already surpass gpt-3.5-turbo-16k’s overall performance on a suite of long-context tasks.
70B LLM Variant Surpasses GPT-3.5 Long-Context Performance
By
–

By
–

2/ Effective Long-Context Scaling with LLMs – propose a 70B variant that can already surpass gpt-3.5-turbo-16k’s overall performance on a suite of long-context tasks.