Robust Critics: Defending LLMs Against Multi-Turn Attacks
让大模型不再“被套路”:多轮攻击通过动态推断用户意图实现防御。
arXiv:2607.20472v1 Announce Type: new Abstract: When a user asks a language model something harmful, is it a genuine attack or a misunderstood but wel…
让大模型不再“被套路”:多轮攻击通过动态推断用户意图实现防御。
arXiv:2607.20472v1 Announce Type: new Abstract: When a user asks a language model something harmful, is it a genuine attack or a misunderstood but wel…
LLM代理正以惊人效率攻破传统网页机器人防御,这篇论文系统揭示了当前安全机制的致命盲区。
arXiv:2607.18659v1 Announce Type: cross Abstract: LLM-based browser agents are rapidly changing the threat landscape for web security. Unlike traditio…
量化部署大模型暗藏后门陷阱?FLipGuard精准防御QCB攻击,守护AI安全底线。
arXiv:2606.28962v1 Announce Type: cross Abstract: Model quantization is essential for the efficient deployment of Large Language Models (LLMs), but in…
被USENIX Security'26收录,系统评估并防御多模态大模型破解验证码的最新研究。
arXiv:2512.02318v4 Announce Type: replace-cross Abstract: This paper studies how multimodal large language models (MLLMs) undermine the security guara…
提出生成语义抗体方法,让视觉-语言模型在开放世界对抗攻击下实现免疫级防御,实验覆盖ImageNet及4个OOD基准。
arXiv:2605.30745v1 Announce Type: new Abstract: Large Vision-Language Models have achieved unprecedented success in zero-shot recognition by aligning …
用Mock工具调用隔离不可信提示输入,为LLM安全防护提供新思路。
arXiv:2605.30521v1 Announce Type: new Abstract: Large language models must frequently process untrusted inputs, such as judging an answer from another…
Web agent面临提示注入威胁,论文提出WARD防御框架,增强对抗鲁棒性,值得安全研究与AI开发者关注。
arXiv:2605.15030v1 Announce Type: cross Abstract: Web agents can autonomously complete online tasks by interacting with websites, but their exposure t…
前沿论文:反蒸馏指纹技术,用于检测LLM被无授权蒸馏,平衡鲁棒性与模型性能。
arXiv:2602.03812v2 Announce Type: replace-cross Abstract: Model distillation enables efficient emulation of frontier large language models (LLMs), cre…