1
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
用机制可解释性方法拆解微调后LLM的道德偏见,揭示潜藏的偏好与风险,为AI对齐提供新视角。
arXiv:2510.12229v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have been shown to internalize human-like biases during finetun…