A recent analysis by OpenAI has uncovered significant flaws within SWE-Bench Pro, a popular coding benchmark for AI models. The findings raise questions about the reliability and accuracy of current methods used to evaluate AI performance in software engineering tasks.
Source: OpenAI