1
LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems
首个聚焦安全关键控制室的LLM操作员多轮红队基准,覆盖对抗鲁棒性与越狱攻击评测。
arXiv:2606.20408v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly proposed as supervisory components for safety-cri…