evaluation-family · Source-linked discovery

Assistant Bench Closed Book One Shot

Tests whether AI agents can perform real-world time-consuming tasks on the web.

Open interactive record →
OriginOri Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya et al.TopicsAutonomy & agentsStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources