1
SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs
破解多模态大模型安全短板,定位并操控通用安全神经元,为AI对齐提供新思路
arXiv:2607.28969v1 Announce Type: new Abstract: Although Large Language Models (LLMs) have demonstrated promising safety performance, extending them t…