I think this makes sense mostly for tasks where you think models have relevant non-shared knowledge. E.g. it’d be good for aidanbench tasks where you want as much coherent diversity as possible — models should be more inspired by other models than by copies of themselves
Using model diversity for better benchmark evaluation
By
–