evaluation-family · Documentary review

Swe Bench

Evaluates AI's ability to resolve genuine software engineering issues sourced from 12 popular Python GitHub repositories, reflecting realistic coding and debugging scenarios.

Open interactive record →
OriginCarlos E. Jimenez, John Yang, Alexander Wettig et al.TopicsGeneral capabilityStatusReviewed

Can support

Patch-generation success on the selected SWE-bench issues under the exact repository, test, scaffold, and execution setup.

Cannot support by itself

General software engineering productivity, requirements discovery, maintainability, security, teamwork, or reliable deployment.

Decision use

Best used for

Comparing coding agents on reproducible repository-level bug-fixing tasks.

Not enough for

Claims that a model can replace software engineers or safely maintain production systems.

Original sources