evaluation-resource · Documentary review

FrontierMath

Expert-written advanced mathematics benchmark intended to remain difficult for frontier models.

Open interactive record →
OriginEpoch AITopicsGeneral capability · AI R&DStatusReviewed

Can support

Success on the evaluated FrontierMath problem set under the stated elicitation, tool, sampling, and grading conditions.

Cannot support by itself

General intelligence, mathematical research autonomy, theorem-proving reliability across the field, AI safety, or AI R&D capability by itself.

Decision use

Best used for

Tracking high-end mathematical problem solving on a difficult, expert-authored and contamination-conscious task set.

Not enough for

Safety conclusions, claims of autonomous mathematical research, or broad comparisons that mix different FrontierMath sets or elicitation budgets.

Original sources