You can see the mean response length for correct and incorrect answers across different models. The source model is on the left and the fine-tuned model on the right. Average performance of the fine-tuned model is reported at the top as "32k score | 4k score".
Fine-tuned Model Response Length Analysis Across AI Models
By
–
