1
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
逐词分解生成过程,新方法轻松突破大模型安全防线,揭示LLM防护盲区。
arXiv:2604.25921v2 Announce Type: replace Abstract: Large Language Models (LLMs) are trained to refuse harmful requests, yet they remain vulnerable to…