AI Dynamics

Global AI News Aggregator

About

Multi-hop Reasoning and Tool Use Benchmark in Language Models

Answering this question correctly is a good mini benchmark challenge, IMO. It requires multi-hop search and reasoning (since the answer isn't directly in the Llama-2 paper, so requires checking its references too) and tools (since it requires doing multi-digit arithmetic).

→ View original post on X — @jeremyphoward