The Power of Backdoor Absorption in Community Training
后门攻击也能被反向吸收?社区训练中的隐蔽威胁与防御新视角,AI安全必读。
arXiv:2607.06643v1 Announce Type: cross Abstract: Backdoor attacks severely threaten large-scale AI models. When model owners delegate training to ext…
后门攻击也能被反向吸收?社区训练中的隐蔽威胁与防御新视角,AI安全必读。
arXiv:2607.06643v1 Announce Type: cross Abstract: Backdoor attacks severely threaten large-scale AI models. When model owners delegate training to ext…
LLM后门检测新方法:通过类子空间正交化解决离散空间触发器逆向难题
arXiv:2606.31309v1 Announce Type: cross Abstract: While post-training backdoor detection and trigger inversion schemes have been developed for AIs use…
本文发现LLM推理中的“悬崖词”——单个token即可导致数学运算失败,揭示模型脆弱性根源。
arXiv:2606.25524v1 Announce Type: new Abstract: Large language models (LLMs) reach high accuracy in mathematical reasoning, but individual traces on t…
多触发木马悄无声息劫持机器人任务规划,AI安全攻防的硬核新研究。
arXiv:2504.17070v3 Announce Type: replace-cross Abstract: Robots need task planning methods to achieve goals that require more than one action. Recent…
提出后门遗忘泛化新路径,让大模型摆脱未知触发器威胁,捍卫LLM安全防线。
arXiv:2606.03785v1 Announce Type: new Abstract: Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can gene…
提出通过分析触发器内部相关性和外部影响来防御GNN后门攻击的新方法,使攻击者陷入两难困境。
arXiv:2605.08278v2 Announce Type: replace-cross Abstract: GNNs have become a standard tool for learning on relational data, yet they remain highly vul…