1
Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect
小语言模型也能内省?本文提出IFT方法,用激活引导训练小模型检测自身扰动,提升可解释性与鲁棒性。
arXiv:2607.14111v1 Announce Type: new Abstract: Can small language models detect and report on perturbations their own internal activations? We invest…