To be persuaded that LLM was reasoning I would want to see (a) an analysis that compared the output with training set in a more serious way than superficial examination of data contamination in the GPT-4 paper & (b) robustness across different formulations of test problems, such
Evaluating LLM Reasoning: Rigorous Analysis Beyond Data Contamination
By
–