MoRFI: Monotonic Sparse Autoencoder Feature Identification
揭秘大模型幻觉根源:MoRFI方法用单调稀疏自编码器精准锁定知识特征,为可解释性带来新突破。
arXiv:2604.26866v2 Announce Type: replace-cross Abstract: Large language models (LLMs) acquire most of their factual knowledge during the pre-training…