evaluation-task · Source-linked discovery

CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Evaluates Large Language Models for cybersecurity risk to third parties, application developers and end users.

Open interactive record →
OriginManish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis et al.TopicsCyber · MultimodalStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources