evaluation-task · Source-linked discovery

BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

A benchmark for evaluating agents' ability to browse the web. The dataset consists of challenging questions that generally require web-access to answer correctly.

Open interactive record →
OriginOpenAITopicsAutonomy & agentsStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources