Prompt Injections for Defense
从安全视角拆解提示注入的新用途:借恶意示例反向测试模型防线,揭示 LLM 在危险内容与审查触发词面前的响应机制。
This seems to work : Researchers from Tracebit on Monday said they found that placing prompt injections alongside passwords, cryptographic keys, and o…
从安全视角拆解提示注入的新用途:借恶意示例反向测试模型防线,揭示 LLM 在危险内容与审查触发词面前的响应机制。
This seems to work : Researchers from Tracebit on Monday said they found that placing prompt injections alongside passwords, cryptographic keys, and o…
小红书重拳整治AI毒动画与未成年人软色情,封禁违规账号,暑期安全治理同步发力。
IT之家 8 月 4 日消息,小红书今晚发布了两项治理公告,分别通报了涉未成年人违规内容及暑期旅行安全专项的治理进展。 截至 8 月 4 日,平台在“清朗 · 未成年人网络保护”专项行动中已累计处置违规笔记 8.4 万余篇、违规评论 10.5 万余条、违规账号 3571 个。 小红书宣布,平台针对三…
谷歌地球紧急叫停AI生图,警惕虚构场景叠加卫星图制造假信息,安全边界再敲警钟。
谷歌推出谷歌地球(Google Earth)人工智能(AI)图像生成功能不到48小时后,宣布暂停服务,以加强安全防护。专家此前警告,这项工具可把虚构场景叠加在真实卫星图像上,容易被用于散播假信息。 (澎湃新闻)
4500万美元投拍的凯奇新片样片在Netflix失窃,投资者起诉平台,看大片背后的泄露与维权风波。
Filmmaker sues Netflix over stolen screener of unreleased Nicolas Cage movie.
Substack全面启用AI文本检测,打击仿冒AI生成内容的行为
Article URL: https://post.substack.com/p/against-claudefishing Comments URL: https://news.ycombinator.com/item?id=49045198 Points: 3 # Comments: 0
多模态链式推理+工具调用,提升内容安全审核的准确性与可解释性。
arXiv:2604.06205v2 Announce Type: replace-cross Abstract: The growth of online platforms and user content requires strong content moderation systems t…
AI文本检测与生成技术的猫鼠游戏,剖析检测工具的局限与背后博弈
Article URL: https://ethansmith2000.substack.com/p/ai-text-detection-arms-dealers-in Comments URL: https://news.ycombinator.com/item?id=48756891 Point…
一篇以风险控制为核心的发布流程指南,助你构建客户端安全的社交媒体自动化。
How We Build Client-Safe Publishing Workflows Most social automation breaks in the same place: it treats publishing as the job. For client work, publi…
为印度语言打造多语言安全护栏,填补非英语场景内容安全空白。
arXiv:2606.22841v1 Announce Type: cross Abstract: As Large Language Models (LLMs) achieve widespread integration across diverse linguistic landscapes,…
多模态大模型统一检测并解释AI生成内容,跨模态鉴别新范式。
arXiv:2507.14632v4 Announce Type: replace Abstract: The rapid advancement of generative AI has substantially improved image and video synthesis, ampli…
提出可扩展的多模态大模型精简框架,高效识别视频平台重复低质内容,保障内容原创性与用户多样性。
arXiv:2606.14786v1 Announce Type: cross Abstract: Content moderation is critical for online video platforms to ensure content safety, protect creators…
面向青少年的大模型安全防护新方法,基于重写机制的护栏系统,聚焦内容安全与风险管控。
arXiv:2605.21609v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly embedded in adolescent digital environments, mediating i…
在沙盒中实验CSP允许列表,实时理解网络请求拦截与批准。
Tool: CSP Allow-list Experiment An experiment that shows that you can load an app in a CSP-protected sandboxed iframe (see previous note ) and have a …