1
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
将论文中的越狱攻击方法转化为可复现的基准测试,填补了LLM安全领域从理论到实践的缺口。
arXiv:2602.24009v4 Announce Type: replace-cross Abstract: Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making …