Reproducing a benchmark is a minor managerial task. It's not even engineering. The scientific hypothesis of LLaMa models is: "If we overtrain the same model architecture on our data, it outperforms previous iteration." Without the data, LLaMa data the experiment is not repro.
LLaMa Reproducibility: Data Access Essential for Scientific Validation
By
–