Best is to evaluate the models with a prompt where the understand the task. If a model is 20% below with a prompt compared to another, clearly you are not evaluating the model properly and it makes no sense to report this number or use it in a comparaison.