evaluation-suite · Source-linked discovery

AgentBoard

Multi-environment benchmark and analysis toolkit for language-model agents.

Open interactive record →
OriginHKUST NLP and the AgentBoard paper authorsTopicsAutonomy & agents · Evaluation integrityStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources