打造 AI 网安“红队”:OpenAI 介绍内部漏洞检测模型 GPT-Red
OpenAI自曝内部“红队”模型GPT-Red,将提示注入攻击成功率从95%压至0.05%,AI安全实战效果惊人。
IT之家 7 月 16 日消息,OpenAI 当地时间 15 日介绍了其内部使用的网络安全“红队”模型 GPT-Red。 该模型可自动化地进行各种网络攻击模拟 ,帮助 OpenAI 提升对外模型产品的鲁棒性。 OpenAI 表示,其过去半年 自 GPT-5.3 后的每个生产模型均将“红队”模型用于训…
OpenAI自曝内部“红队”模型GPT-Red,将提示注入攻击成功率从95%压至0.05%,AI安全实战效果惊人。
IT之家 7 月 16 日消息,OpenAI 当地时间 15 日介绍了其内部使用的网络安全“红队”模型 GPT-Red。 该模型可自动化地进行各种网络攻击模拟 ,帮助 OpenAI 提升对外模型产品的鲁棒性。 OpenAI 表示,其过去半年 自 GPT-5.3 后的每个生产模型均将“红队”模型用于训…
首个面向多智能体AI系统的提示注入基准测试工具,填补安全评估空白
Article URL: https://freyzo.github.io/deep-xpia/ Comments URL: https://news.ycombinator.com/item?id=48549498 Points: 1 # Comments: 0
LLM防护盾反成攻击靶心:揭示基于大模型的agent安全护栏存在拒绝服务新漏洞
arXiv:2606.14517v1 Announce Type: cross Abstract: LLM-based guardrails have emerged as a highly effective defense against prompt injection and jailbre…
脑机接口+LLM代理面临新型注入攻击,这个研究揭示了神经授权通道的隐蔽风险。
arXiv:2606.09315v1 Announce Type: cross Abstract: BCI-to-agent pipelines turn decoded neural activity into an authorization channel for tool-use agent…
通过对抗日志内容对LLM安全运营系统发起提示注入攻击,揭示新型安全威胁。
arXiv:2605.24421v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as analyst assistants in security operations cent…
揭秘多智能体LLM系统中域伪装注入攻击如何绕过安全防护,揭示现有检测器的盲区。
arXiv:2605.22001v1 Announce Type: cross Abstract: Injection detectors deployed to protect LLM agents are calibrated on static, template-based payloads…