The interesting question is why and when would this work? Intuitively this looks like variance reduction in a setup where variance dominates bias because the LLM has very high capacity. So consensus makes sense. Thoughts @roydanroy @ArnaudDoucet1 ?? It has vibes of online