Has anyone tried RL to rewrite prompts for reasoning models to further improve outputs? I'm assuming so, it feels pretty obvious, but if not I want to try it. If you know of any existing work here, pls lmk so I don't re-do something people have already done!
Using RL to Optimize Prompts for Reasoning Models
By
–