ToxScreen: Detecting Whether an LLM Has Been Poisoned
针对大模型投毒攻击的检测新方法ToxScreen,通过分析模型行为判断是否被恶意篡改
arXiv:2607.26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training…
针对大模型投毒攻击的检测新方法ToxScreen,通过分析模型行为判断是否被恶意篡改
arXiv:2607.26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training…
稀疏自编码器新应用:高效防御LLM越狱攻击,来自ICML可解释性研讨会的前沿研究。
arXiv:2602.12418v2 Announce Type: replace-cross Abstract: Jailbreak attacks remain a persistent threat to large language model safety. We propose Cont…
最新研究从理论上证明,任何通用防护措施都无法完全阻止大模型越狱攻击,揭示了LLM安全的根本困境
Article URL: https://github.com/brandoncarl/llm-jailbreaking/blob/main/On%20the%20Impossibility%20of%20Perfect%20Universal%20Guardians%20Against%20LLM…
用输出重写破坏多轮越狱攻击,语义保持下高效防御,AI安全新方法。
arXiv:2606.02640v1 Announce Type: cross Abstract: Multi-turn jailbreak attacks pose a growing threat to large language model (LLM) safety because they…
提出内部化逐步反思机制,让模型自主识别并防御间接越狱攻击,AI安全新范式。
arXiv:2605.20654v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sop…