RLHF sacrifices in-context learning for ease of prompting. Model is more capable in an informal, less quantifiable sense. Some contrived tasks are nearly impossible without tuning though e.g. summarizing text that’s close in length to the size of the context window.
RLHF Trade-offs: In-Context Learning vs. Prompting Ease in LLMs
By
–