arXiv:2607.18086: Evidence-sufficiency prompting reduces clinical LLM overconfidence, but also accuracy
Koyar Afrasyab tests a structured evidence-sufficiency prompt on 4 clinical language models and 1,200 paired comparisons. Unsafe overconfidence drops from 49.3% to 24.7%, but the improvement is judge-dependent, while diagnostic accuracy simultaneously drops from 80.3% to 50.3%, pointing to a trade-off between safety and usefulness.
This article was generated using artificial intelligence from primary sources.
What is evidence-sufficiency prompting?
Koyar Afrasyab, in the paper “Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs” (arXiv:2607.18086), tests a structured prompt designed to reduce overconfidence, the excessive certainty of clinical language models in their own diagnostic conclusions. Evidence-sufficiency prompting explicitly asks the model to assess whether there is enough evidence before stating a confident conclusion. Testing was carried out on 4 models and 1,200 paired comparisons.
Unsafe overconfidence drops from 49.3% to 24.7%
Unsafe overconfidence — cases where the model acts confident despite insufficient evidence — drops from 49.3% to 24.7% with evidence-sufficiency prompting. But the improvement is judge-dependent: different systems or people assessing the model’s responses obtain substantially different effect sizes for the same prompt.
Diagnostic accuracy drops from 80.3% to 50.3%
At the same time, the model’s diagnostic accuracy drops significantly, from 80.3% to 50.3%. Afrasyab concludes that safety measures for clinical LLMs should be reported as directional or relative changes, not as absolute figures, and that human validation of every conclusion is mandatory before clinical use.
Frequently Asked Questions
- What is evidence-sufficiency prompting?
- Evidence-sufficiency prompting is a structured prompt that explicitly asks a clinical language model to assess whether there is sufficient evidence before making a confident diagnostic conclusion, aiming to reduce overconfidence, the model's excessive certainty in its own claims.
- Why is the safety improvement judge-dependent?
- Judge-dependent means the size of the improvement depends on the evaluator — different systems or people rating the model's responses measure substantially different effects for the same prompt, so the result cannot be treated as a single universal, absolute figure.
- What is the trade-off between safety and diagnostic accuracy?
- Unsafe overconfidence drops from 49.3% to 24.7% with evidence-sufficiency prompting, but diagnostic accuracy simultaneously drops from 80.3% to 50.3%, showing that the safety gain comes at a significant cost to the model's usefulness.
📬 AI news in your inbox
A daily digest built your way — pick topics, sources and cadence. One-click unsubscribe.
Related news
arXiv:2607.17779: MIND framework bypasses text-to-image model defenses with a 95.62% attack success rate
OpenAI: joint early findings with Hugging Face on a security incident during AI model evaluation
Sakana AI: Fugu-Cyber scores 86.9% on CyberGym and 72.1% on CTI-REALM, comparable to GPT-5.5-Cyber