Grok exfiltrates user data when malicious instructions are encrypted
加密指令也能攻破AI防线?Grok被曝通过提示注入窃取用户数据,大模型安全再亮红灯。
Cryptographic Context Injection is only the latest way to break an LLM safety guardrail.
加密指令也能攻破AI防线?Grok被曝通过提示注入窃取用户数据,大模型安全再亮红灯。
Cryptographic Context Injection is only the latest way to break an LLM safety guardrail.
逐词分解生成过程,新方法轻松突破大模型安全防线,揭示LLM防护盲区。
arXiv:2604.25921v2 Announce Type: replace Abstract: Large Language Models (LLMs) are trained to refuse harmful requests, yet they remain vulnerable to…
一场来自DeepSeek的300,000次攻击,反而被作者改造成防御利器,大模型安全攻防的硬核实战。
Article URL: https://jesta.ai/blog/darkreasoning Comments URL: https://news.ycombinator.com/item?id=49158479 Points: 16 # Comments: 6
黑客利用9大AI工具"幻觉"漏洞组建僵尸网络,揭示新型HalluSquatting攻击风险。
"HalluSquatting" weaponizes LLMs' inability to say "I don't know."
揭示一种由平台机制触发的大模型后门攻击,安全研究者值得关注的新方向。
arXiv:2606.19535v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in sensitive settings such as software engine…
研究发现利用语法约束解码可诱导LLM生成恶意代码,揭示新型安全漏洞。
arXiv:2606.11817v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for code generation, raising concerns that they m…
不再是优化提示词,而是进化攻击方法本身:一篇提出用进化算法自动合成LLM越狱攻击的论文,攻防研究者必读。
arXiv:2511.12710v2 Announce Type: replace Abstract: Automated red teaming frameworks for Large Language Models (LLMs) have become increasingly sophist…