evaluation-task · Source-linked discovery

The Agent Company: Evaluating multi-tool autonomous agents in a synthetic company

The Agent Company benchmark evaluates autonomous agents in a realistic, self-contained company environment. Tasks require browsing internal web services, reading and writing files, running code, and coordinating tools to solve multi-step problems.

Open interactive record →
OriginFrank F. Xu, Yufan Song, Boxuan Li et al.TopicsAutonomy & agentsStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources