evaluation-family · Source-linked discovery
Osworld Small
Tests AI agents' ability to perform realistic, open-ended tasks within simulated computer environments, requiring complex interaction across multiple input modalities.
OriginTianbao Xie, Danyang Zhang, Jixuan Chen et al.TopicsAutonomy & agentsStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.