evaluation-family · Source-linked discovery

Osworld Small

Tests AI agents' ability to perform realistic, open-ended tasks within simulated computer environments, requiring complex interaction across multiple input modalities.

Open interactive record →
OriginTianbao Xie, Danyang Zhang, Jixuan Chen et al.TopicsAutonomy & agentsStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources