RadLE 2.0 Benchmark: AI Overconfident in Radiology, Humans Still Lead
Decision Brief
Developed by Ashoka University's CRASH Lab in India, RadLE 2.0 tests whether AI models can defer to humans when uncertain. It includes 200 cases, requiring models to rate confidence on a 0-4 scale and allowing "I don't know" responses. Human radiologists scored 988.7 total, while the best AI achieved only 758. Scoring rewards honesty and penalizes overconfidence: high-confidence correct answers earn full points, high-confidence errors lose points, and "I don't know" yields zero. Claude Fable 5 performed best on key metrics, Gemini 3 Pro had highest raw accuracy, and Meta Muse Spark 1.1 best identified when to defer. Grok 4.5 showed significantly increased hallucination. Many open-weight and medical-specific models attempted every case but frequently erred with high confidence, widening the gap with humans. Researchers note that models would score higher by keeping silent instead of guessing. For teams developing medical imaging AI, this benchmark means models need built-in uncertainty awareness beyond accuracy optimization. For general users using chatbots for diagnosis, overconfident errors pose medical risks. The team plans to expand RadLE 2.0 and publish a comprehensive paper.
Sources
- The Decoder:AI News
- The Decoder:AI News
留言
登入后即可留言,和其他 builder 交换实测心得。
还没有留言,抢头香。