Automating Deception: Scalable Multi-Turn LLM Jailbreaks
大模型安全新漏洞:多轮对话如何用“登门槛”心理绕过对齐机制
arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-…
大模型安全新漏洞:多轮对话如何用“登门槛”心理绕过对齐机制
arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-…
让大模型不再“被套路”:多轮攻击通过动态推断用户意图实现防御。
arXiv:2607.20472v1 Announce Type: new Abstract: When a user asks a language model something harmful, is it a genuine attack or a misunderstood but wel…
多轮对话破解LLM防线,分解信用分配实现高效越狱攻击,揭示AI安全新漏洞
arXiv:2607.11070v1 Announce Type: new Abstract: Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jail…
多轮多模态攻击防不胜防?这篇论文提出预测性防御机制,应对未见过的新型攻击。
arXiv:2605.18988v1 Announce Type: cross Abstract: The expansion of Multimodal Large Language Models (MLLMs) and their integration into autonomous agen…