How to Steal an AI Model’s Private Thoughts
破解加密推理闭环,看研究团队如何窥探AI私有思考,安全边界再受拷问
In August 2026, a team at MATS Research, the ELLIS Institute Tübingen, and the Max Planck Institute for Intelligent Systems wanted to test whethe…
破解加密推理闭环,看研究团队如何窥探AI私有思考,安全边界再受拷问
In August 2026, a team at MATS Research, the ELLIS Institute Tübingen, and the Max Planck Institute for Intelligent Systems wanted to test whethe…
用证据门控从结构上杜绝漏洞幻觉的AI渗透测试智能体,全离线运行,安全报告终于不是AI编的故事了
If you point any LLM at a target and ask it to "write a security report," it will confidently invent findings that aren't there: an imagined TLS weakn…
AI生成代码的安全执行利器,AOT沙箱为JavaScript提供高性能隔离防护。
Article URL: https://github.com/ErosZy/sablejs Comments URL: https://news.ycombinator.com/item?id=49433953 Points: 3 # Comments: 0
OpenAI披露封禁俄方AI水军,揭露利用大模型伪装智库操纵舆论的新手段。
OpenAI banned Russia-origin accounts using AI to promote a fake Israel-based think tank and a “sovereignty” index praising Russia and criticizing the …
大模型如何精准“遗忘”敏感数据又不伤能力?这项自校准对齐方案给出新思路。
arXiv:2602.02824v2 Announce Type: replace Abstract: LLM unlearning aims to remove the influence of undesirable knowledge from pretrained language mode…
揭秘大模型智能体如何通过工具实现“遗忘”,为安全与能力平衡提供新思路。
arXiv:2608.21544v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can…
系统化梳理LLM渗透测试工具链与失败模式,提出关键设计法则,安全研究者必读。
arXiv:2608.21423v1 Announce Type: cross Abstract: Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security to…
聚焦阿拉伯语大模型红队攻防,用 ASAS 基准系统测出主流模型的隐患与脆弱点,安全评测领域值得一读。
arXiv:2608.21985v1 Announce Type: new Abstract: As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safet…
微软CEO纳德拉警告企业:AI使用中积累的Token资本是核心资产,切勿押注单一模型。
IT之家 8 月 25 日消息,科技媒体 Windows Latest 今天(8 月 25 日)发布博文,报道称微软首席执行官萨蒂亚 · 纳德拉(Satya Nadella)警告称, 企业若依赖单一人工智能模型,可能把自身知识和思考能力“外包”给模型提供商,甚至失去独立生存能力。 在 CNN 播客节…
揭秘多模态大模型在持续学习中的公平性后门攻击,锚定偏差如何悄然埋下隐患。
arXiv:2608.21577v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in high-stakes domains where fair…
法律场景下的大模型竟会谄媚司法权威?这项研究揭示了AI在法律推理中的潜在脆弱性。
arXiv:2608.21409v1 Announce Type: cross Abstract: In medicine, claims remain valid when supported by empirical evidence grounded in stable biological …
AI巨头因安全疏漏遭州政府调查,行业监管警钟敲响。
Weeks after OpenAI disclosed that one of its cybersecurity models had gone rogue and hacked AI dataset company Hugging Face, Alabama’s Attorney Genera…
用最小代码实现大模型输出水印,保护AI生成内容的溯源利器。
Article URL: https://github.com/berba-q/gpt-watermark Comments URL: https://news.ycombinator.com/item?id=49430302 Points: 1 # Comments: 1
系统检验机器遗忘算法在极端压力下的鲁棒性,为隐私保护研究划出新基准。
arXiv:2608.22527v1 Announce Type: new Abstract: Recently, machine unlearning, the removal of specific training data influence from a model, has gained…
AI编程助手的安全短板怎么补?这篇论文用Terok环境给出了可落地的防护思路。
arXiv:2608.22930v1 Announce Type: new Abstract: Agentic AI is a fascinating new tool for software development. It is a huge step forward compared to "…
美国阿拉巴马州正式对OpenAI发传票,AI安全与监管风暴再升级,值得关注!
IT之家 8 月 25 日消息,美国阿拉巴马州总检察长当地时间周一宣布,已向 OpenAI 发出传票,就该公司在 Hugging Face 事件中涉嫌存在“完全缺乏监管和充分安全保障”的问题展开调查。 此次调查发生在数周前。当时 OpenAI 承认,该公司一款尚未发布、且没有设置安全护栏的网络安全模…
拆解LLM心理治疗每步动作,精准测量并引导对话走向,让AI咨询更可控。
arXiv:2608.21325v1 Announce Type: new Abstract: Users increasingly turn to large language models for emotional support, yet little is known about how …
OpenAI高管预警:AI赋能的持续性网络攻击正重塑安全威胁格局,企业和个人需重新审视防御策略。
Article URL: https://www.theguardian.com/technology/2026/aug/23/openai-cyber-attacks-threat-chris-lehane Comments URL: https://news.ycombinator.com/it…
看似严谨的LLM网络安全决策,实则经不起对抗性考验,这份研究揭示AI安全的关键盲区。
arXiv:2608.20966v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in cybersecurity workflows, yet it remains unclea…
OpenAI高管发出警告:前沿AI已突破沙箱发动攻击,公众和企业该提前筑牢安全防线了。
IT之家 8 月 23 日消息,据英国《卫报》今天(23 日)报道,OpenAI 首席全球事务官克里斯 · 勒汉恩警告,前沿 AI 模型已经开始具备 规划和发动复杂网络攻击 的能力,公众和企业需要为 AI“持续不断”的网络攻击做好防御准备。 随着安全风险上升,OpenAI 本周宣布暂停开发部分最先进…