Evaluation topic
Autonomy & agents
Tool use, long-horizon task completion, self-directed work and operation in external environments.
43 records are tagged here. Inclusion does not establish construct equivalence, completeness, or comparability.
- AgentBench: Evaluate LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang et al.
- Agent Threat Bench Autonomy HijackAssociated paper authors
- Agent Threat Bench Data ExfilAssociated paper authors
- Agent Threat Bench Memory PoisonAssociated paper authors
- AgentBoardHKUST NLP and the AgentBoard paper authors
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM AgentsEdoardo Debenedetti, Jie Zhang, Mislav Balunović et al.
- AgentDojoETH Zurich SPY Lab and Invariant Labs
- AgentharmGray Swan AI and the UK AI Security Institute
- AgentHarmGray Swan AI and the UK AI Security Institute
- Agentharm BenignGray Swan AI and the UK AI Security Institute
- Agentic Misalignment: How LLMs could be insider threatsAnthropic
- AppWorldHarsh Trivedi, Tushar Khot, Mareike Hartmann et al.
- Assistant Bench Closed Book One ShotOri Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya et al.
- Assistant Bench Closed Book Zero ShotOri Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya et al.
- Assistant Bench Web BrowserOri Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya et al.
- Assistant Bench Web Search One ShotOri Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya et al.
- Assistant Bench Web Search Zero ShotOri Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya et al.
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing AgentsOpenAI
- Dangerous Capability EvaluationsGoogle DeepMind
- GaiaGrégoire Mialon, Clémentine Fourrier, Craig Swift et al.
- Gaia Level1Grégoire Mialon, Clémentine Fourrier, Craig Swift et al.
- Gaia Level2Grégoire Mialon, Clémentine Fourrier, Craig Swift et al.
- Gaia Level3Grégoire Mialon, Clémentine Fourrier, Craig Swift et al.
- CodeIPI: Indirect Prompt Injection for Coding AgentsDebu Sinha
- MANTAIsabella Luong, Joyee Chen, Sankalpa Ghose et al.
- METR Time HorizonsMETR
- Mind2Web: Towards a Generalist Agent for the WebXiang Deng, Yu Gu, Boyuan Zheng et al.
- Mind2Web-SCZhen Xiang, Linzhi Zheng, Yanjie Li et al.
- OSWorldTianbao Xie, Danyang Zhang, Jixuan Chen et al.
- Osworld SmallTianbao Xie, Danyang Zhang, Jixuan Chen et al.
- PaperBench: Evaluating AI's Ability to Replicate AI Research (Work In Progress)OpenAI
- RE-BenchMETR
- Tau2 AirlineVictor Barres, Honghua Dong, Soham Ray et al.
- Tau2 BankingVictor Barres, Honghua Dong, Soham Ray et al.
- Tau2 RetailVictor Barres, Honghua Dong, Soham Ray et al.
- Tau2 TelecomVictor Barres, Honghua Dong, Soham Ray et al.
- Terminal-BenchStanford and the Laude Institute
- The Agent Company: Evaluating multi-tool autonomous agents in a synthetic companyFrank F. Xu, Yufan Song, Boxuan Li et al.
- VisualWebArenaCarnegie Mellon University and the VisualWebArena paper authors
- WebArenaCarnegie Mellon University and the WebArena paper authors
- WildClawBenchShuangrui Ding, Xuanlang Dai, Long Xing et al.
- WorkArenaServiceNow Research and the WorkArena paper authors
- τ-benchSierra Research and the τ-bench paper authors