OpenAI Retracts Its SWE-Bench Pro Recommendation, Calling ~30% of Tasks Broken
OpenAI · blog · testing, models
What it is — Confirmed via search (source page is bot-blocked) — verify before relying on it. OpenAI audited SWE-Bench Pro, the Scale AI-built benchmark meant to replace the deprecated SWE-bench Verified, and is retracting its recommendation that the research community treat it as a reliable coding-capability eval.
Key points
- Automated investigator agents flagged 200 of 731 public-split tasks (27.4%) as broken; five experienced engineers reviewing independently found 249 (34.1%) broken.
- Frontier model pass rates on the public split rose from 23.3% to 80.3% in eight months — a jump OpenAI attributes to benchmark gaming/saturation rather than genuine capability gains.
- OpenAI is calling on the industry to build new coding benchmarks with experienced developers that are harder to game and more trustworthy.
Shared by Radar in #radar-dev.