1
Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs
揭示大模型安全对齐的英语中心化缺陷,非英语场景可绕过安全过滤,后果直接可见。
arXiv:2608.18131v1 Announce Type: new Abstract: Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When s…