1
Explaining Intrinsic Moral Self-Correction with Mechanistic Interpretability
机械可解释性揭示大模型内在道德自我修正的机制,为AI安全对齐提供新视角。
arXiv:2505.11924v4 Announce Type: replace-cross Abstract: Intrinsic moral self-correction refers to the phenomenon where a language model refines its …