Dynamic Jailbreaking Attack
打破静态优化桎梏,动态越狱攻击同时提升大模型攻击的有效性、效率与灵活性,安全研究者需警惕新范式。
arXiv:2510.02422v4 Announce Type: replace-cross Abstract: Existing gradient-based jailbreak attacks typically optimize a fixed-length adversarial suff…
打破静态优化桎梏,动态越狱攻击同时提升大模型攻击的有效性、效率与灵活性,安全研究者需警惕新范式。
arXiv:2510.02422v4 Announce Type: replace-cross Abstract: Existing gradient-based jailbreak attacks typically optimize a fixed-length adversarial suff…
探讨破解蒸馏防御的真正含义,揭示对抗样本攻击与防御博弈的最新深度研究。
arXiv:2606.25059v1 Announce Type: cross Abstract: Black-box LLMs (accessible only via API) are vulnerable to distillation attacks, in which an attacke…
揭示“想象-行动”世界模型的致命漏洞,用预言机级攻击篡改AI对未来的信任,安全研究者必读。
arXiv:2606.22966v1 Announce Type: cross Abstract: Many recent vision-language-action (VLA) policies adopt an imagine-then-act design. A world-action m…
一篇揭示RAG-LM安全训练中“注入悖论”的论文:注入逆向抑制品牌推荐,带来新的安全思考。
arXiv:2606.09204v1 Announce Type: new Abstract: We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injec…
提出Stable-GFlowNet,用对比轨迹平衡生成多样对抗样本,提升LLM红队测试的鲁棒性
arXiv:2605.00553v2 Announce Type: replace Abstract: Large Language Model (LLM) Red-Teaming, which proactively identifies vulnerabilities of LLMs, is a…
教你用数据投毒对抗AI:让模型失效的实战策略与工具解析
Article URL: https://www.youtube.com/watch?v=Z8aLGHmnRyc Comments URL: https://news.ycombinator.com/item?id=48321906 Points: 2 # Comments: 0