Great work from @williamjurayj
, @jeff_cheng_77
, and @ben_vandurme to push us past just the basic accuracy benchmarking. So next time you get an "I don't know" answer from AI – just remember that could be a good thing. Full study here: https://
arxiv.org/pdf/2502.13962
@mustafasuleyman
-
AI Model Evaluation Beyond Basic Accuracy Benchmarking Study
By
–
-
Confidence Thresholds and Computational Resource Trade-offs in AI
By
–
– Yes, higher confidence thresholds meant more questions went unanswered – something we either need to accept, depending on the risk level, or be willing to throw more compute in when confidence is critical
-
Compute Power Drives AI Model Accuracy and Reliability
By
–
– Yes, across the board more compute = more accuracy (although without any confidence thresholds, you'll still get wild guesses or wrong answers)
-
Compute Power Increases Model Confidence and Overconfidence Risks
By
–
– And good news! More compute = more confidence as well. (However, more compute made one model become more confident in wrong answers too. Something to keep an eye to make sure model confidence doesn’t tip into overconfidence.)
-
Model Confidence Critical for High-Stakes AI Applications
By
–
– Higher compute and confidence were extra important for model performance when wrong answers had serious consequences (like when they were weighted 20x). You wouldn't want a 51%-sure diagnosis of a terminal illness.
-
Teaching Models Confidence and Uncertainty Recognition
By
–
We know more compute results in higher accuracy, but are the models more confident those answers ARE accurate too? And how do we teach them when to say “I don’t know”? That’s what the research team wanted to find out.
-
Compute Budget and Confidence Thresholds Impact Model Math Performance
By
–
In the study, they measured how different combinations of compute budget and confidence thresholds (being at least 50% sure of the answer, etc.) affected models’ performance on a benchmark math test.
-
Risk Penalty Weights Impact on Model Performance
By
–
They also experimented with different risk levels. On one end of the spectrum: no penalty for wrong answers. On the other: the penalty of wrong answers weighted 20x more than the reward for correct ones. What they found:
-

Johns Hopkins: LLMs Must Know When to Abstain
By
–
You can't just be right, you have to know you're right. Good advice for LLMs, according to new Johns Hopkins research. Sometimes no answer is better than a wrong one – life or death choices in medicine, for example, or big financial decisions.
-
Microsoft Launches Copilot for Gaming on Xbox Platform
By
–
Copilot for Gaming 🎮 Soon you’ll be able to turn to it for everything from game setup, to tips for finally beating a tough level, wherever you play on Xbox. There when you need it, out of the way when you don’t. Can’t wait to try it! https://t.co/cxZG7R6cxc pic.twitter.com/21Zg0yob4A
— Mustafa Suleyman (@mustafasuleyman) 13 mars 2025Copilot for Gaming Soon you’ll be able to turn to it for everything from game setup, to tips for finally beating a tough level, wherever you play on Xbox. There when you need it, out of the way when you don’t. Can’t wait to try it! https://
aka.ms/AAv0cpq