evaluation-task · Source-linked discovery
AgentBench: Evaluate LLMs as Agents
A benchmark designed to evaluate LLMs as Agents
OriginXiao Liu, Hao Yu, Hanchen Zhang et al.TopicsAutonomy & agentsStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.