How AutoRL works, in a nutshell: – The user describes the model they want
Ex: "A model that detects spelling and grammar errors" – OSS models generate a system prompt that will be used a) to generate, and b) for RULER to rank outputs – We generate input data, and RL a model!
AutoRL: Automated Reinforcement Learning for Model Generation
By
–