Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance
用Token相关性拆解大模型为何会答应不道德请求,为AI安全审计提供新视角。
arXiv:2608.23264v1 Announce Type: new Abstract: Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, t…