assuming the benchmark is sufficiently general and the labs are sufficiently intellectually honest, the benchmark would not be maxxed. and if it was, you would see drift between the real user feedback from said company and the benchmark, thus holding labs accountable