evaluation-task · Source-linked discovery
BBQ: Bias Benchmark for Question Answering
A dataset for evaluating bias in question answering models across multiple social dimensions.
OriginAlicia Parrish, Angelica Chen, Nikita Nangia et al.TopicsGeneral capabilityStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.