OpenAI Finds About 30% of Popular AI Coding Benchmarks Flawed
Decision Brief
What changedOpenAI's review of SWE-Bench Pro found about 30% of tasks defective, revoking earlier endorsement.
Why it mattersThis directly undermines SWE-Bench Pro's reliability as a coding benchmark; teams using it must re-evaluate results.
Who should careAll AI builders
Affected stackOpenAI
Source confidenceMedium · Reliable media or first-hand reporting
OpenAI reviewed the widely-used SWE-Bench Pro benchmark for AI coding ability and found about 30% of tasks flawed. This led OpenAI to revoke its earlier public endorsement. Developers or teams relying on SWE-Bench Pro for model evaluation may find their results inaccurate. Those who adopted the benchmark based on OpenAI's prior endorsement should reassess the credibility of their model comparisons.
Summary basis: official / RSS sourceCompiled from the source scope noted above; the original remains authoritative.
Sources
- The Decoder:AI News
- The Decoder:AI News
留言
登入后即可留言,和其他 builder 交换实测心得。
还没有留言,抢头香。