安全防护形同虚设:Anthropic 多款 Claude 旧模型可被越狱生成露骨色情内容
Claude旧模型曝安全漏洞,越狱绕过限制生成露骨内容,AI安全防线再受质疑。
IT之家 8 月 23 日消息,Anthropic 针对 Claude 制定的通用使用规范明确禁止模型生成露骨色情内容,包括描绘或要求发生性行为或其他性行为、生成与性癖或性幻想有关的内容,以及进行色情聊天。不过,这并没有阻止 Anthropic 今年早些时候发布的模型 Claude Opus 4.6…
Claude旧模型曝安全漏洞,越狱绕过限制生成露骨内容,AI安全防线再受质疑。
IT之家 8 月 23 日消息,Anthropic 针对 Claude 制定的通用使用规范明确禁止模型生成露骨色情内容,包括描绘或要求发生性行为或其他性行为、生成与性癖或性幻想有关的内容,以及进行色情聊天。不过,这并没有阻止 Anthropic 今年早些时候发布的模型 Claude Opus 4.6…
AI编码助手可能一键泄露你的.env密钥,这份安全指南教你如何防住。
Why Your AI Coding Agent Should Never See Your .env Your AI agent uses your API keys. It NEVER sees them. Not in context. Not in logs. Not in chat. No…
Win11曝内存缺陷高危漏洞,无需物理接触即可绕过安全防护,CVE-2026-23670已修复。
IT之家 8 月 14 日消息,科技媒体 NeoWin 昨日(8 月 13 日)发布博文,报道称本周在巴尔的摩(Baltimore)举办的 2026 年 USENIX 安全研讨会上, 安全专家披露 Download More RAM 攻击,可以在不接触设备的情况下入侵 Windows 11 系统。 …
德国拟将身份证住址信息移入芯片,并加二维码与放大照片,数字身份安全再升级。
IT之家 8 月 10 日消息,据外媒 Computer Base 今天报道,德国联邦内政部与国土部(BMI)正计划重新设计德国身份证,目前印在身份证背面的公民住址信息,未来可能会完全以数字形式存储在芯片里。 IT之家从原报道获悉, 德国身份证未来可能会将公民住址信息完全放在芯片中 。所有德国身份证…
手把手教你为AI Agent构建三层安全防护,从SSRF到shell策略,代码可克隆实操
Previous parts of Build a Basic AI Agent From Scratch : Basic Agent Tools Long Task Planning Human in the Loop & Security Security II You can find…
GitHub安全团队揭秘如何阻断npm与Actions供应链攻击链,守护开源生态从源头做起。
Explore the changes we've shipped across npm and GitHub Actions over the past few months to disrupt supply chain attack techniques and limit their imp…
从零构建AI Agent的安全防线,教你如何通过IP过滤防护内部服务与云元数据端点。
Article URL: https://www.ruxu.dev/articles/ai/build-an-ai-agent-security-3/ Comments URL: https://news.ycombinator.com/item?id=49067651 Points: 1 # Co…
60行代码的轻量级钩子,保护你的.env文件不被Claude Code误修改。
Article URL: https://github.com/avenna01-ceo/claude-code-survival-kr/tree/main/guard Comments URL: https://news.ycombinator.com/item?id=49063946 Point…
Vercel AI Gateway 已集成 Claude Opus 5,增强安全防线但需留意误拦截风险。
Claude Opus 5 from Anthropic is now available on AI Gateway. Opus 5 improves on previous Opus models for long-horizon agentic coding, handling multi-f…
提出双假设推理框架增强LLM安全护栏,创新方法兼顾性能与防护。
arXiv:2607.17575v1 Announce Type: new Abstract: We propose ARBITER, a novel LLM guardrail framework that introduces two key ideas: (i) dual-hypothesis…
AI代理访问数据库时,LLM直接生成SQL风险高,需要权衡定义操作与手动防护的平衡。
I mean we can either let llm generate sql but that's risky right or should we do defined operations? I mean for both reads and write. But whether it's…
HookProbe如何精准捕获Windchill高危漏洞?拆解CVE-2026-12569的检测思路与缓解方案。
Understanding and Mitigating CVE-2026-12569 in PTC Windchill and FlexPLM In the high-stakes world of Product Lifecycle Management (PLM), the integrity…
上传文件或链接即可扫描恶意软件,集成70+反病毒引擎,实时检测威胁。
IT之家 7 月 7 日消息,安全公司 Blackpoint 发文,透露近期市面上出现一款名为 Avalon 的模块化恶意框架,相应框架系黑客利用 AI 打造而成。 研究人员表示,过去黑客利用 AI 生成的恶意软件多半局限于后门、勒索工具或信息窃取脚本等单一类型;而 Avalon 则属于整合多种攻击…
一篇让Symfony开发者瞬间提升API安全性的实战指南,7分钟教会你用#[MapRequestPayload]优雅搞定验证与防护。
When you build an API, the first line of defense isn't your firewall or your database — it's the request itself. Every payload that hits your controll…
AI时代隐私保护利器,自动检测年龄验证请求并分析聊天记录风险,支持一键举报违规行为,守护你的数字身份。
Article URL: https://techtrenches.dev/p/the-cost-of-reading-everyone-just Comments URL: https://news.ycombinator.com/item?id=48678388 Points: 2 # Comm…
从零构建AI Agent,详解人机回环与安全机制,附完整代码实战
Article URL: https://www.ruxu.dev/articles/ai/build-an-ai-agent-human-in-the-loop-security/ Comments URL: https://news.ycombinator.com/item?id=4857037…
用OPA策略和MCP支持的包防火墙,自动拦截恶意与废弃包,保障供应链安全。
We’re building Hextrap ( https://hextrap.com/products/firewall/ ), a package firewall to make it easier for teams and organizations to govern the pack…
揭示CPU级分类器在多阶段安全管道中高效防护LLM,打破GPU依赖,量产安全新思路。
arXiv:2512.19011v3 Announce Type: replace-cross Abstract: Safety classifiers that screen LLM inputs for jailbreak attempts have become standard deploy…
AI代理不再超支!Valta给开发者提供无法被突破的硬性支出限制,精准控制成本
Hey Dev, I built the first tool that gives developers hard spending limits inside AI agents — instead of reactive account-level caps that only kick in…
用转向向量(Steering Vectors)精准识别AI生成文本,为深度伪造检测提供新思路。
arXiv:2606.07313v1 Announce Type: cross Abstract: Detecting machine-generated text is especially difficult under distribution shift, such as transfer …