1
Learning from the Self-future: On-policy Self-distillation for dLLMs
探索将自蒸馏技术从自回归模型迁移至扩散LLM,突破传统方法限制,为后训练提供新思路。
arXiv:2606.18195v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) has proven effective for post-training large language models (LLMs)…