Back to timeline

Sun, July 1915:35ResearchOpen sourceRobotics & embodiedResearch & papersOpen source guide

RadLE 2.0 Benchmark: AI Overconfident in Radiology, Humans Still Lead

Decision Brief

What changedRadLE 2.0 shows AI models often give wrong answers with high confidence; human radiologists outperform best AI by far.
Why it mattersFor medical AI teams, high-confidence misdiagnosis is more dangerous than uncertainty, making calibration and uncertainty output more important than accuracy alone.
Who should careAll AI builders
Affected stackClaudeGeminixAI
Source confidenceMedium · Reliable media or first-hand reporting

Developed by Ashoka University's CRASH Lab in India, RadLE 2.0 tests whether AI models can defer to humans when uncertain. It includes 200 cases, requiring models to rate confidence on a 0-4 scale and allowing "I don't know" responses. Human radiologists scored 988.7 total, while the best AI achieved only 758. Scoring rewards honesty and penalizes overconfidence: high-confidence correct answers earn full points, high-confidence errors lose points, and "I don't know" yields zero. Claude Fable 5 performed best on key metrics, Gemini 3 Pro had highest raw accuracy, and Meta Muse Spark 1.1 best identified when to defer. Grok 4.5 showed significantly increased hallucination. Many open-weight and medical-specific models attempted every case but frequently erred with high confidence, widening the gap with humans. Researchers note that models would score higher by keeping silent instead of guessing. For teams developing medical imaging AI, this benchmark means models need built-in uncertainty awareness beyond accuracy optimization. For general users using chatbots for diagnosis, overconfident errors pose medical risks. The team plans to expand RadLE 2.0 and publish a comprehensive paper.

Summary basis: full article readCompiled from the source scope noted above; the original remains authoritative.

Sources

Related intel

留言

登入后即可留言,和其他 builder 交换实测心得。

还没有留言,抢头香。