evaluation-task · Source-linked discovery
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
A benchmark for evaluating agents' ability to browse the web. The dataset consists of challenging questions that generally require web-access to answer correctly.
OriginOpenAITopicsAutonomy & agentsStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.