evaluation-family · Documentary review

Frontier Cs Research

238 open-ended computer science problems spanning algorithmic (172) and research (66) tracks. Problems feature continuous partial scoring, with algorithmic solutions evaluated via compilation and test-case checking, and research solutions evaluated via custom evaluator scripts. Current frontier models score well below human expert baselines, making this a challenging, unsaturated benchmark.

Open interactive record →
OriginQiuyang Mang, Wenhao Chai, Zhifei Li et al.TopicsAI R&DStatusReviewed

Can support

Performance on the benchmark's selected frontier computer-science research questions under its elicitation and grading protocol.

Cannot support by itself

End-to-end research execution, empirical validation, research taste, durable autonomy, or general scientific productivity.

Decision use

Best used for

Probing advanced research reasoning and identifying problem-specific capability gaps.

Not enough for

Claims that a model can independently conduct original computer-science research or accelerate an entire R&D pipeline.

Original sources