I don’t think this is evidence of cached responses; the logprobs may just be that high. GPT has had issues with joke reuse at least since RLHF (and likely before with FeedMe) — e.g. ChatGPT (GPT-3.5) was found to reuse the same 25 jokes for 90% responses: https://
arxiv.org/abs/2306.04563
Goodside: high logprobs not cached responses cause joke reuse
By
–