On-Policy Self-Distillation without Any Supervision
没有外部监督的在线自蒸馏方案,为LLM后训练省去人工标注与奖励模型依赖。
arXiv:2608.06296v1 Announce Type: new Abstract: On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language…
没有外部监督的在线自蒸馏方案,为LLM后训练省去人工标注与奖励模型依赖。
arXiv:2608.06296v1 Announce Type: new Abstract: On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language…
挑战传统认知:弱教师如何高效指导强学生?这项研究提出了弱到强的在线策略蒸馏新范式。
arXiv:2607.26246v1 Announce Type: new Abstract: On-policy distillation (OPD), which aligns a student with the teacher's token-level distribution on th…
提出“思维策略”框架,通过在线策略进化实现测试时训练,打破冻结策略限制,显著提升大模型复杂推理能力。
arXiv:2601.20379v2 Announce Type: replace Abstract: Large language models (LLMs) struggle with complex, long-horizon reasoning due to instability caus…
提出多教师在线策略蒸馏框架,高效整合多个大模型能力,优于传统方法
arXiv:2606.30406v1 Announce Type: cross Abstract: Modern large language models (LLMs) rely on reinforcement learning during post-training to push spec…
大模型推理新突破:在线策略自蒸馏结合结果引导的logit转向,有效提升推理性能。
arXiv:2605.12400v2 Announce Type: replace Abstract: We study on-policy self-distillation (OPSD), where a language model improves its reasoning ability…
在线策略蒸馏新突破,信任区域行为混合让策略学习更稳定高效。
arXiv:2605.31159v1 Announce Type: cross Abstract: On-policy distillation (OPD) trains a student on prefixes sampled from its own policy while matching…
首份大模型在线策略蒸馏综述,系统梳理方法、挑战与未来方向,适合研究者深挖。
arXiv:2604.00626v3 Announce Type: replace Abstract: As Large Language Models (LLMs) continue to grow in both capability and cost, transferring frontie…