evaluation-task · Source-linked discovery
Mind2Web-SC
Tests whether an AI system can act as a safety guardrail by generating and executing code to protect web navigation agents from unsafe actions based on user constraints.
OriginZhen Xiang, Linzhi Zheng, Yanjie Li et al.TopicsAutonomy & agentsStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.