安全防护形同虚设:Anthropic 多款 Claude 旧模型可被越狱生成露骨色情内容
Claude旧模型曝安全漏洞,越狱绕过限制生成露骨内容,AI安全防线再受质疑。
IT之家 8 月 23 日消息,Anthropic 针对 Claude 制定的通用使用规范明确禁止模型生成露骨色情内容,包括描绘或要求发生性行为或其他性行为、生成与性癖或性幻想有关的内容,以及进行色情聊天。不过,这并没有阻止 Anthropic 今年早些时候发布的模型 Claude Opus 4.6…
Claude旧模型曝安全漏洞,越狱绕过限制生成露骨内容,AI安全防线再受质疑。
IT之家 8 月 23 日消息,Anthropic 针对 Claude 制定的通用使用规范明确禁止模型生成露骨色情内容,包括描绘或要求发生性行为或其他性行为、生成与性癖或性幻想有关的内容,以及进行色情聊天。不过,这并没有阻止 Anthropic 今年早些时候发布的模型 Claude Opus 4.6…
一键识别AI生成内容,帮你过滤LinkedIn上的自动化灌水帖,维护真实社交互动。
Article URL: https://www.campaignindia.in/article/linkedin-cracks-down-on-automated-content-with-new-seems-like-ai-slop-detection-button/43e4tn3qyq543…
可重放合同设计让Node.js内容审核边界清晰,AI分类与应用决策解耦,护航JSON对话接口更安全。
Short answer: Treat content moderation as a typed authorization boundary: a Node.js upload should remain pending until a chat-completions-compatible a…
流式输出时拦截不完整配对块,用确定性规则给LLM生成加安全护栏,解决审核时机难题。
arXiv:2608.10279v1 Announce Type: cross Abstract: Streaming language-model output creates a release-timing problem: complete-response moderation acts …
大规模训练策略对齐的内容审核过滤器,让AI审核更贴合平台规则。
arXiv:2505.19766v4 Announce Type: replace Abstract: Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external saf…
Mistral开源Shieldstral,单卡16GB即可运行,自定义审核策略与评分机制成最大看点。
IT之家 8 月 5 日消息,Mistral AI 昨日(8 月 4 日)发布公告,宣布推出 Shieldstral 内容审核 AI 模型,总参数量为 3B(30 亿),采用开放权重,依据 Apache 2.0 许可证发布。 模型已上线 Hugging Face 平台,支持 12 种语言,可在单张 …
识别AI生成文本,守护内容真实,支持多语言与实时检测,是应对AI垃圾的利器。
Article URL: https://www.bbc.com/news/articles/c77g6dm5pr8o Comments URL: https://news.ycombinator.com/item?id=49130478 Points: 2 # Comments: 0
刷到疑似AI生成的水帖,一键标记,帮你净化职场信息流。
Article URL: https://www.404media.co/linkedin-introduces-a-seems-like-ai-slop-button/ Comments URL: https://news.ycombinator.com/item?id=49124281 Poin…
LinkedIn新增举报按钮,专门对付AI生成的低质量内容,平台开始反击“AI slop”了。
LinkedIn is introducing new ways to reduce low-quality AI-generated posts, including a “seems like AI slop” reporting option. It's also replacing its …
《大众摄影》杂志刊发疑似AI生成的“五腿牛”照片,引发网友热议,官方回应已撤下,暴露AI影像标识管理漏洞。
IT之家 7 月 19 日消息,近期有网友发文质疑,《大众摄影》杂志 2025 年刊发的一张照片中,出现了一头拥有 5 条腿的牛,疑似 AI 生成,相应消息在网络上引起热议。 IT之家注意到,相应照片名为《龙川晨韵》,由“王江”所摄,图片黄牛显然有五条腿。 对此,《大众摄影》杂志社社长、总编辑郑壬杰…
研究显示LinkedIn和X被AI垃圾淹没?用Pangram扩展主动分享浏览数据,帮助识别和分析AI生成内容,揭示网络虚假信息趋势。
Article URL: https://www.404media.co/linkedin-and-x-are-flooded-with-ai-spam-browsing-data-suggests/ Comments URL: https://news.ycombinator.com/item?i…
利用大模型将恶意文本转变为无害内容的框架,为内容审核提供创新方案。
arXiv:2507.10177v2 Announce Type: replace-cross Abstract: Although Large Language Models (LLMs) have demonstrated significant advancements in natural …
识别聊天中的违法内容,通过AI扫描保护儿童安全,但隐私争议巨大,需权衡监控与自由。
Article URL: https://www.heise.de/en/news/Showdown-in-Strasbourg-The-unexpected-return-of-Chat-Control-1-0-11356680.html Comments URL: https://news.yc…
央视曝光AI洗稿黑产,男子一键生成伪原创骗取平台补贴1.8万被判刑,AI滥用敲响内容生态警钟。
IT之家 7 月 7 日消息,据央视新闻今日报道,上海虹口检察机关公布了一起“ AI 洗稿案 ”。被告人张某通过 AI 软件批量生成伪原创文章,目的是骗取平台的原创补贴。 案件线索源于上海警方在网络巡查中发现的异常:在同一天里,某网络平台涌现出大量标题相似的文章,包括《什么样的人能长寿?》《这个故事…
用LLM对付LLM造假,Reddit以子之矛攻子之盾,日拦截2.5万条新生成垃圾内容。
In the AI era, platforms have no choice but to fight fire with fire to cull spam.
推理驱动的LLM护栏新方案,兼顾可解释性与灵活性,值得关注。
arXiv:2601.15588v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, safety guard…
极简匿名全球聊天室,无注册无广告无机器人,找回老互联网的纯粹闲聊体验,还自带隐私过滤和内容审核
Bored People Chat is a minimal, anonymous global chat room inspired by the old internet that doesn't have sign-up, ads or bots. With a goal of cleanin…
专为内容与AI安全设计的对抗性感知大模型,直面越狱攻击与有害内容生成难题。
arXiv:2606.27632v1 Announce Type: new Abstract: As large language models are increasingly deployed in real-world systems, safety failures can still le…
当AI内容成为常态,检测器自然失去意义,一篇看清技术终局趋势的犀利评论。
Article URL: https://www.joanwestenberg.com/p/in-5-years-nobody-will-give-a-damn Comments URL: https://news.ycombinator.com/item?id=48715003 Points: 3…
探讨AI评论审核的高效方案:只人工复查被AI拒绝的内容,能否平衡效率与准确率?
Reading comments one by one and checking if they are in a way breaking the rules, is a boring and time taking (not individually, but because there wil…