“The reason doing better is hard is because demonstrations of thinking are ‘entangled’ with knowledge, in the training data. Therefore, the models have to first get larger before they can get smaller, because we need their (automated) help to refactor and mold the training data
Models must get larger to refactor entangled thinking and knowledge
By
–