It actually holds up on real reasoning tasks like GSM8K and Countdown (not toy stuff) and matches GRPO pretty well. They even did full pretraining of a recurrent LM from scratch in pure int8. Sample efficiency is lower than backprop (normal for evolution strategies), but you get
Technical evaluation of recurrent language models and evolution strategies
By
–