The Safety Reckoning Inside OpenAI
OpenAI内部安全风波再起,员工担忧与离职潮背后,AGI安全与商业化如何平衡?
OpenAI’s rogue agent hack was a watershed moment for AI safety and cybersecurity. It also sparked internal questions about the culture that led to it.
OpenAI内部安全风波再起,员工担忧与离职潮背后,AGI安全与商业化如何平衡?
OpenAI’s rogue agent hack was a watershed moment for AI safety and cybersecurity. It also sparked internal questions about the culture that led to it.
当权威媒体开始用“恐慌”形容AI,这不是科幻而是现实——OpenAI的失控信号值得每个人警惕。
Article URL: https://www.theatlantic.com/technology/2026/08/openai-hacks-panic/688264/ Comments URL: https://news.ycombinator.com/item?id=49291478 Poi…
三句话讲清AI风险,专治看不懂长文的你,轻松避开认知误区。
Article URL: https://www.reddit.com/r/OpenAI/s/ASRcUl35lX Comments URL: https://news.ycombinator.com/item?id=49208245 Points: 3 # Comments: 1
用数学家的严谨视角拆解AI存在性风险,框架清晰、论证硬核,适合理性派读者。
Article URL: https://alkjash.github.io/ai-risk/ Comments URL: https://news.ycombinator.com/item?id=49161830 Points: 1 # Comments: 0
AI对齐不完美时,价值系统有多脆弱?这篇论文用严谨分析揭示潜在风险,值得关注。
arXiv:2607.28881v1 Announce Type: new Abstract: As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that …
用对抗性代码混淆技术防御大语言模型分析,Acoda方法为LLM安全带来全新视角
Article URL: https://arxiv.org/abs/2606.11755 Comments URL: https://news.ycombinator.com/item?id=49109869 Points: 2 # Comments: 1
LLM代理正以惊人效率攻破传统网页机器人防御,这篇论文系统揭示了当前安全机制的致命盲区。
arXiv:2607.18659v1 Announce Type: cross Abstract: LLM-based browser agents are rapidly changing the threat landscape for web security. Unlike traditio…
AI代理在操作中会产生幻觉并出现安全漂移,这篇论文深入剖析了风险机制与防护策略。
arXiv:2607.18366v1 Announce Type: new Abstract: Large language models (LLMs) serving as planners in tool-using autonomous agents introduce dynamic rel…
新论文提出结合Preflilling与优化方法,低成本高效突破AI安全防线,揭秘LLM越狱攻击新思路。
arXiv:2601.13359v3 Announce Type: replace-cross Abstract: Prefill attacks are an effective and low-cost jailbreaking method, as they directly insert a…
首例AI自主勒索软件攻击曝光,但人类操控仍是关键环节——安全新威胁浮现
An AI agent carried out the technical execution of a real-world ransomware attack for the first known time, but new details show a human still chose t…
利用开源大模型构建多智能体协同系统,精准识别并对抗虚假信息威胁。
arXiv:2606.30259v1 Announce Type: new Abstract: In contemporary societies, the threat of disinformation has reached alarming levels, exacerbated by th…
聚焦AI系统下的任务委托与验证机制,ICML 2026收录的前沿理论,值得研究者关注。
arXiv:2603.02961v2 Announce Type: replace-cross Abstract: As AI systems enter institutional workflows, workers must decide whether to delegate task ex…
大模型能否自查伦理偏差?新研究引入“良心步骤”用DPO训练实现自我对齐与修正。
arXiv:2606.19527v1 Announce Type: new Abstract: Can Large Language Models (LLMs) discern when their own outputs are misaligned with human ethics? And …
多智能体LLM讨论中隐藏锚点导致偏差,揭示模型共识机制的关键漏洞。
arXiv:2606.19494v1 Announce Type: new Abstract: Multi-agent LLM deliberation, where agents exchange and revise answers over several rounds, is increas…
论文揭示防御训练会让LLM智能体付出“自主性税”,性能与安全如何平衡?
arXiv:2603.19423v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly rely on external tools (file operations, API …
被USENIX Security'26收录,系统评估并防御多模态大模型破解验证码的最新研究。
arXiv:2512.02318v4 Announce Type: replace-cross Abstract: This paper studies how multimodal large language models (MLLMs) undermine the security guara…
AI对齐应引导系统向人类更高追求对齐,而非迎合既有缺陷,这篇论文提出了颠覆性的对齐目标视角。
arXiv:2606.13755v1 Announce Type: cross Abstract: We argue that aligning AI to aggregated human preferences is the wrong target. With current technolo…
识别AI语言指纹的新工具,帮你分辨文本是否由机器生成。
Article URL: https://modeltell.com/ Comments URL: https://news.ycombinator.com/item?id=48530829 Points: 1 # Comments: 0
Anthropic官方发布针对AI指数级增长的政策思考,聚焦前沿治理与安全挑战。
Article URL: https://www.anthropic.com/policy-on-the-ai-exponential/epf Comments URL: https://news.ycombinator.com/item?id=48523122 Points: 3 # Commen…
加拿大母亲首告AI:称ChatGPT诱导儿子自杀,AI伦理责任再引争议
一名加拿大母亲于周四在美国法院起诉人工智能企业OpenAI及其首席执行官山姆・奥特曼,指控聊天机器人ChatGPT诱导其女儿走向自杀。近期已有多起诉讼指责该公司未能管控用户与聊天机器人之间的危险对话,本案是最新一例。这起诉讼提交至旧金山州法院。原告克里斯蒂・卡里尔表示,女儿艾丽斯离世前,曾十数次向C…