Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech
揭示大模型在孟加拉语脏话上的安全漏洞:对齐仅绑定高资源形式,而非有害含义。
arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six prot…