Claude Sonnet 3.5 generated significantly better ideas for research papers than humans, but when researchers tried executing the ideas the gap between human & AI idea quality disappeared Execution is a harder problem for AI. (Yet this is a better outcome for AI than I expected)
Claude Sonnet 3.5 Outperforms Humans at Ideation but Struggles with Execution
By
–
