Suspecting court of using AI, man injected prompts in filings to try to win case
专治AI被"下蛊"的注入攻击检测器!自动模拟恶意指令轰炸AI系统,扫描越狱漏洞生成防御报告,防止被套路
Judge warns pro se litigants are using chatbots wrong and getting desperate.
专治AI被"下蛊"的注入攻击检测器!自动模拟恶意指令轰炸AI系统,扫描越狱漏洞生成防御报告,防止被套路
Judge warns pro se litigants are using chatbots wrong and getting desperate.
基于论文提出的工作流级越狱方法,检测IDE编码代理在任务分解、文件生成等环节的安全漏洞,适合安全研究者评估AI助手
arXiv:2607.03968v1 Announce Type: cross Abstract: Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks…
利用中间层熵的动态变化检测大模型越狱,为AI安全提供新思路。
arXiv:2606.25182v1 Announce Type: cross Abstract: Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted p…
跨语言越狱检测新突破:不依赖语言特征,学习意图表示,有效识别多种语言下的越狱攻击。
arXiv:2606.11202v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in applications for global multilingual users, …
浅层神经网络集成策略GuardNet,专攻大模型提示注入与越狱攻击的鲁棒检测
arXiv:2606.05566v1 Announce Type: new Abstract: Large Language Models (LLMs) have transformed natural language processing, but they remain vulnerable …
解决LLM安全评估的两大缺陷:多数据集统一阈值与隐含操作点透明化,这套评估框架更严谨。
arXiv:2606.02959v1 Announce Type: new Abstract: Published evaluations of prompt-injection and jailbreak detectors for Large Language Models often suff…
提出基于锚定token级logits的LLM越狱检测方法SelfGrader,高效识别对抗性攻击。
arXiv:2604.01473v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are powerful tools for answering user queries, yet they remain …
高效多任务安全分类器,低成本实时检测毒性、越狱、仇恨言论等,专为LLM安全防护设计。
arXiv:2605.29659v1 Announce Type: cross Abstract: Real-time safety filtering for large language model (LLM) applications requires classifiers that can…