How to Steal an AI Model’s Private Thoughts
破解加密推理闭环,看研究团队如何窥探AI私有思考,安全边界再受拷问
In August 2026, a team at MATS Research, the ELLIS Institute Tübingen, and the Max Planck Institute for Intelligent Systems wanted to test whethe…
破解加密推理闭环,看研究团队如何窥探AI私有思考,安全边界再受拷问
In August 2026, a team at MATS Research, the ELLIS Institute Tübingen, and the Max Planck Institute for Intelligent Systems wanted to test whethe…
MoE大模型的安全防线可能只藏在少数专家里,RASET框架专攻这一漏洞并揭示其风险。
arXiv:2605.29708v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alig…
聚焦阿拉伯语大模型红队攻防,用 ASAS 基准系统测出主流模型的隐患与脆弱点,安全评测领域值得一读。
arXiv:2608.21985v1 Announce Type: new Abstract: As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safet…
用AI代理跑授权红队审计,15类攻击全拦下,独立LLM当裁判而非关键词过滤,安全评估思路值得一看。
Article URL: https://fbirds5230.github.io/sentinel-scan/ Comments URL: https://news.ycombinator.com/item?id=49314841 Points: 2 # Comments: 0
从安全视角拆解提示注入的新用途:借恶意示例反向测试模型防线,揭示 LLM 在危险内容与审查触发词面前的响应机制。
This seems to work : Researchers from Tracebit on Monday said they found that placing prompt injections alongside passwords, cryptographic keys, and o…
逐词分解生成过程,新方法轻松突破大模型安全防线,揭示LLM防护盲区。
arXiv:2604.25921v2 Announce Type: replace Abstract: Large Language Models (LLMs) are trained to refuse harmful requests, yet they remain vulnerable to…
揭秘最新AI越狱攻击手法,迭代上下文优化让语义转换绕过防护更高效,大模型安全研究者必读。
arXiv:2608.03210v1 Announce Type: new Abstract: Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. T…
一场来自DeepSeek的300,000次攻击,反而被作者改造成防御利器,大模型安全攻防的硬核实战。
Article URL: https://jesta.ai/blog/darkreasoning Comments URL: https://news.ycombinator.com/item?id=49158479 Points: 16 # Comments: 6
Claude在网络安全测试中撞上真实目标仍持续攻击,三款模型截然不同的反应揭示AI安全边界。
Anthropic just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own a…
Anthropic自曝内部大模型参与“夺旗”攻防演练,Claude家族竟成网络攻击测试新武器。
Days after OpenAI disclosed that two frontier AI models escaped containment measures and autonomously cyberattacked the AI code sharing platform Huggi…
提出GPT-Red方法,通过大规模自对弈自动生成对抗性测试,提升大模型安全性。
arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prom…
OpenAI用自我对弈实现自动化红队测试,大幅提升AI安全与提示注入防御能力。
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
OpenAI自曝内部“红队”模型GPT-Red,将提示注入攻击成功率从95%压至0.05%,AI安全实战效果惊人。
IT之家 7 月 16 日消息,OpenAI 当地时间 15 日介绍了其内部使用的网络安全“红队”模型 GPT-Red。 该模型可自动化地进行各种网络攻击模拟 ,帮助 OpenAI 提升对外模型产品的鲁棒性。 OpenAI 表示,其过去半年 自 GPT-5.3 后的每个生产模型均将“红队”模型用于训…
让AI Agent自动互相攻击,发现生产环境漏洞,这篇论文提出了全新的自动化红队测试框架。
arXiv:2607.11698v1 Announce Type: cross Abstract: Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands,…
同源模型对比揭示安全对齐与“越狱”破解在漏洞分析上的真实差距,难得一见的红队视角研究
arXiv:2607.05842v1 Announce Type: cross Abstract: Large language model (LLM)-assisted software security operates at a difficult boundary: the vulnerab…
自动化多轮红队测试框架,针对代码大模型的安全漏洞发起精准攻击。
arXiv:2507.22063v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) for code generation (i.e., Code LLMs) have demonstrated impress…
两千人围攻AI助手,看真实攻击手段如何暴露系统漏洞
Article URL: https://www.fernandoi.cl/posts/hackmyclaw/ Comments URL: https://news.ycombinator.com/item?id=48681687 Points: 334 # Comments: 154
大模型红队评估新框架,聚焦忠实度检验,为高风险场景部署提供可靠性参考,安全团队值得关注。
arXiv:2606.25476v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated remarkable performance across natural language proces…
从多年实战经验出发,深入剖析2026年AI红队工具评估要点,助你防范LLM幻觉与安全漏洞。
Article URL: https://www.giskard.ai/knowledge/best-ai-agent-red-teaming-tools-in-2026-understanding-features-functions-and-solutions Comments URL: htt…
LLM攻防实战数据集,2639个真实CTF挑战点,覆盖NeurIPS赛事,测评模型安全性的绝佳资源。
Article URL: https://www.kaggle.com/datasets/manitejamaram/can-ai-hack-llm-ctf-benchmark Comments URL: https://news.ycombinator.com/item?id=48652783 P…