Thought about this more: It doesn’t make sense it’s the prompt, because inference costs. Still feels human-written — but why would you RLHF into reciting all policy verbatim? ChatGPT falls for the same trick and the text it recites is just parameters.
Skepticism about prompt vs human-written policy in ChatGPT RLHF
By
–