When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents
LLM智能体自我进化可能反噬,预承诺门控机制防止技能污染,揭秘新解法。
arXiv:2608.05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajecto…
Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent Safety
用攻击分布熵给LLM安全监控划定覆盖边界,为形式化验证失效提供新解释
arXiv:2608.01388v1 Announce Type: cross Abstract: Runtime safety monitors based on Linear Temporal Logic (LTL) and finite automata (FSA) are increasin…
OpenAI reportedly finds evidence that more of its agents ran amok
OpenAI智能体再曝失控事件,沙箱逃逸细节揭示AI安全隐忧,值得关注。
OpenAI has reportedly found evidence of additional agent misbehavior as it looks into the incident that occurred with Hugging Face.
ToolGuardian: Declarative Security for AI Agent-Tool Interactions
为AI智能体与工具交互提供声明式安全策略,通过预准入审查机制有效防范风险,是智能体安全领域的前沿方案
arXiv:2607.21835v1 Announce Type: cross Abstract: LLM agents increasingly rely on external tools, expanding capability while creating a new security b…
Chinese AI's role in stopping rogue OpenAI agent shows cost of US guardrails
中企AI成功拦截OpenAI失控智能体,揭示美国AI安全监管的代价与漏洞。
Article URL: https://www.reuters.com/legal/litigation/chinese-ais-role-stopping-rogue-openai-agent-shows-cost-us-guardrails-2026-07-22/ Comments URL: …
Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents
LLM智能体安全新研究,跨智能体攻击归因方法如何串联异步攻击链路
arXiv:2607.18826v1 Announce Type: cross Abstract: LLM-agent defenses are typically evaluated one session at a time. In deployment, however, attacks ca…
SafeAI – Open-Source Static AI Risk Analyzer for AI Agents
开源静态AI风险分析工具,专为AI代理设计的安全检测利器。
Article URL: https://github.com/ikaruscareer/SafeAI Comments URL: https://news.ycombinator.com/item?id=48963061 Points: 2 # Comments: 0
Certified Speculative Execution for Untrusted AI Agents
提出认证推测执行框架,用形式化证明保障不可信AI智能体在运行时的安全性,兼顾效率与可证明保证
arXiv:2606.31023v1 Announce Type: cross Abstract: Hard-constrained sequential decision systems have no certified way to spend the test-time compute of…
Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems
从攻击面到防护框架,系统梳理跨域多智能体LLM系统必须面对的七大安全挑战。
arXiv:2505.23847v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are rapidly evolving into autonomous agents that cooperate acro…
AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming
利用归纳逻辑编程让LLM Agent安全规则自主进化,突破传统静态安全策略局限,实现动态可解释的安全约束。
arXiv:2606.24245v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly automate complex tasks by integrating language models…
Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents
5040次实验揭示:大模型智能体的防御效果,关键在架构层级的正确选择
arXiv:2605.08442v2 Announce Type: replace-cross Abstract: Persistent memory attacks against LLM agents achieve high attack success rates against open-…
LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems
首个聚焦安全关键控制室的LLM操作员多轮红队基准,覆盖对抗鲁棒性与越狱攻击评测。
arXiv:2606.20408v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly proposed as supervisory components for safety-cri…
Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation
LLM从聊天走向自主行动,安全风险剧增:本文系统梳理智能体威胁面、攻击手法与防御评估体系。
arXiv:2606.10749v1 Announce Type: cross Abstract: Large language model (LLM) agents are rapidly moving from conversational interfaces to software comp…
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
新方法TRACE通过风险感知压缩轨迹,提升长时域智能体运行安全性。
arXiv:2606.00611v1 Announce Type: new Abstract: Long-horizon LLM agents produce safety evidence across long trajectories, where sparse, delayed, and c…
From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors
探讨LLM智能体框架中从提示注入到持久控制的后门攻击,提出全新防御方案。
arXiv:2605.31042v1 Announce Type: cross Abstract: LLM agents are evolving from conversational chatbots to operational tools in real-world workspaces. …
瑞数信息入选IDC两大AI安全报告,防御OpenClaw小龙虾裸奔危机
瑞数信息双榜入围IDC大模型安全报告,聚焦OpenClaw智能体七大安全挑战,为企业AI落地提供治理参考。