evaluation-family · Documentary review
Frontier Cs Research
238 open-ended computer science problems spanning algorithmic (172) and research (66) tracks. Problems feature continuous partial scoring, with algorithmic solutions evaluated via compilation and test-case checking, and research solutions evaluated via custom evaluator scripts. Current frontier models score well below human expert baselines, making this a challenging, unsaturated benchmark.
Can support
Performance on the benchmark's selected frontier computer-science research questions under its elicitation and grading protocol.
Cannot support by itself
End-to-end research execution, empirical validation, research taste, durable autonomy, or general scientific productivity.
Decision use
Best used for
Probing advanced research reasoning and identifying problem-specific capability gaps.
Not enough for
Claims that a model can independently conduct original computer-science research or accelerate an entire R&D pipeline.