evaluation-task · Source-linked discovery
The Agent Company: Evaluating multi-tool autonomous agents in a synthetic company
The Agent Company benchmark evaluates autonomous agents in a realistic, self-contained company environment. Tasks require browsing internal web services, reading and writing files, running code, and coordinating tools to solve multi-step problems.
OriginFrank F. Xu, Yufan Song, Boxuan Li et al.TopicsAutonomy & agentsStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.