evaluation-family · Documentary review
Swe Bench
Evaluates AI's ability to resolve genuine software engineering issues sourced from 12 popular Python GitHub repositories, reflecting realistic coding and debugging scenarios.
OriginCarlos E. Jimenez, John Yang, Alexander Wettig et al.TopicsGeneral capabilityStatusReviewed
Can support
Patch-generation success on the selected SWE-bench issues under the exact repository, test, scaffold, and execution setup.
Cannot support by itself
General software engineering productivity, requirements discovery, maintainability, security, teamwork, or reliable deployment.
Decision use
Best used for
Comparing coding agents on reproducible repository-level bug-fixing tasks.
Not enough for
Claims that a model can replace software engineers or safely maintain production systems.