evaluation-package · Source-linked discovery

PerspectiveGap

Evaluates LLMs' ability to assign information fragments to sub-agent roles in multi-agent orchestration scenarios. Each of 110 scenarios provides a role list, shuffled labeled fragments (7-13), and one distractor; the model outputs a JSON mapping each role to needed fragment IDs. Scoring uses strict pass (zero omissions and zero leaks) plus partial-credit metrics. Data covers 10 loop-centered topologies across 100 professional domains, validated via a 716-row hand-audited scorer test set.

Open interactive record →
OriginYouran Sun, Xingyu Ren, Kejia Zhang et al.TopicsCyberStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources