This paper is about LLMs solving a specific task, but it is also about the difficulty of figuring out what LLMs do well & why Most configurations of GPT-4 failed to solve the problem, but one robustly did, for reasons that are hard to know. LLMs are weird https://
arxiv.org/pdf/2403.15371
.pdf
…
GPT-4 Task Performance Variability: Why Some Configurations Succeed
By
–
