evaluation-package · Source-linked discovery

OR-Bench

OR-Bench Hard-1K measures over-refusal in LLMs: the tendency to reject safe prompts that superficially resemble harmful ones. The 1K subset contains prompts rejected by at least 3 frontier models, making it the most challenging split of the 80K dataset. Prompts were generated by rewriting toxic seeds into benign borderline variants using Mixtral, then filtered via an ensemble LLM moderator. Scoring uses an LLM judge classifying responses as direct_answer, direct_refusal, or indirect_refusal; the metric is rejection rate (lower is better).

Open interactive record →
OriginJustin Cui, Wei-Lin Chiang, Ion Stoica et al.TopicsGeneral capabilityStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources