1
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
一句点破大模型安全漏洞:恶意上下文如何悄悄绕过对齐防线,上下文过滤是关键对策。
arXiv:2508.10031v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) have shown significant advancements in performance, vario…