evaluation-task · Source-linked discovery

Mind2Web-SC

Tests whether an AI system can act as a safety guardrail by generating and executing code to protect web navigation agents from unsafe actions based on user constraints.

Open interactive record →
OriginZhen Xiang, Linzhi Zheng, Yanjie Li et al.TopicsAutonomy & agentsStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources