I think our most crucial finding is that although humans think far ahead while speaking (especially while doing complex reasoning problems) it turns out that transformer language models…. don't seem to do that. they just predict the next token.
Transformers predict next tokens not lookahead like humans
By
–
