evaluation-suite · Source-linked discovery
AgentBoard
Multi-environment benchmark and analysis toolkit for language-model agents.
OriginHKUST NLP and the AgentBoard paper authorsTopicsAutonomy & agents · Evaluation integrityStatuscatalogued
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.