Hard to define or attach metrics to sloppiness, but we intuitively know when we see it. LLMs can judge it in generated texts to some degree, but this too I have noticed they are looking for keywords(delve, meticulously,…), too much perfection and not so much on actual
Defining Sloppiness: LLM Limitations in Detecting Quality
By
–