evaluation-task · Source-linked discovery

AgentBench: Evaluate LLMs as Agents

A benchmark designed to evaluate LLMs as Agents

Open interactive record →
OriginXiao Liu, Hao Yu, Hanchen Zhang et al.TopicsAutonomy & agentsStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources