GPT-Red: Automated Red Teaming via Self-Play at Scale
提出GPT-Red方法,通过大规模自对弈自动生成对抗性测试,提升大模型安全性。
arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prom…
提出GPT-Red方法,通过大规模自对弈自动生成对抗性测试,提升大模型安全性。
arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prom…
AI安全里程碑:GPT-RED发现人类从未见过的「假思维链」攻击,AI攻防进入自我博弈时代。
OpenAI发布新模型GPT-Red,美国热泵使用率上升,科技前沿速览。
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. …
OpenAI用自我对弈实现自动化红队测试,大幅提升AI安全与提示注入防御能力。
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
OpenAI自曝内部“红队”模型GPT-Red,将提示注入攻击成功率从95%压至0.05%,AI安全实战效果惊人。
IT之家 7 月 16 日消息,OpenAI 当地时间 15 日介绍了其内部使用的网络安全“红队”模型 GPT-Red。 该模型可自动化地进行各种网络攻击模拟 ,帮助 OpenAI 提升对外模型产品的鲁棒性。 OpenAI 表示,其过去半年 自 GPT-5.3 后的每个生产模型均将“红队”模型用于训…
OpenAI打造LLM超级黑客GPT-Red,与GPT-5.6对抗训练,让模型安全防御能力大幅提升。
OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberatta…