AI RADARМоделі
239•07/09/2026, 11:15•3 min read
OpenAI: SWE-Bench Pro is no longer suitable for model evaluation
#OpenAI#benchmarks#AI models#SWE-Bench Pro

OpenAI has announced that SWE-Bench Pro has serious issues that undermine its effectiveness. They are calling for the development of new benchmarks for model testing.
AI Radar • Plus Tier
Full Engineering Breakdown Available in Plus Plan
Daily technical analyses, benchmark reviews, model changes, and vibe coding case studies unlock in the Plus plan.
Discuss in community
Share your questions and insights with developers