SelFusion: Self-distillation for Diffusion Language Models
扩散语言模型的自蒸馏新方法登上ACL 2026,生成效率与质量或迎来新突破。
arXiv:2608.22898v1 Announce Type: new Abstract: Diffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) larg…
扩散语言模型的自蒸馏新方法登上ACL 2026,生成效率与质量或迎来新突破。
arXiv:2608.22898v1 Announce Type: new Abstract: Diffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) larg…
单次推理评估多细则易受干扰,新方法用自蒸馏提升LLM裁判准确率与效率
arXiv:2608.14684v1 Announce Type: cross Abstract: LLM judges increasingly evaluate responses against fine-grained rubric checklists. When a sample req…
没有外部监督的在线自蒸馏方案,为LLM后训练省去人工标注与奖励模型依赖。
arXiv:2608.06296v1 Announce Type: new Abstract: On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language…
用自蒸馏技术增强基于评分标准的强化学习,提升训练效果和泛化能力。
arXiv:2607.18082v1 Announce Type: cross Abstract: Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognize…
扩散语言模型遇冷?dOPSD让它在自蒸馏中越学越强,生成质量突破新高度。
arXiv:2607.04428v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, of…
无需人工标注,通过神经元激活模式筛选数据,实现LLM高效自蒸馏训练。
arXiv:2607.02460v1 Announce Type: cross Abstract: Post-training large language models (LLMs) without real-world interaction feedback or human-labeled …
自蒸馏新方法,让模型在保持思考能力的同时完成策略蒸馏,值得AI研究者一读。
arXiv:2607.02234v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, wh…
用模型自己的输出来训练自己,简单自蒸馏让代码生成能力显著提升,这招值得所有LLM开发者关注。
arXiv:2604.01193v2 Announce Type: replace Abstract: Can a large language model (LLM) improve at code generation using only its own raw outputs, withou…
探索将自蒸馏技术从自回归模型迁移至扩散LLM,突破传统方法限制,为后训练提供新思路。
arXiv:2606.18195v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) has proven effective for post-training large language models (LLMs)…
无需保留集!新方法SHRED用自蒸馏+logit降级实现LLM高效遗忘,拒绝灾难性性能下降。
arXiv:2605.07482v2 Announce Type: replace-cross Abstract: Machine unlearning for large language models (LLMs) aims to selectively remove memorized con…
自蒸馏让大模型在难题上学会专家推理,摆脱依赖更强模型或采样正确解的局限
arXiv:2602.02405v2 Announce Type: replace-cross Abstract: Improving the reasoning capabilities of large language models (LLMs) typically relies either…
一种新颖的反思式策略自蒸馏方法,让语言模型通过自我反思实现跨领域推理能力提升,无需人工标注。
arXiv:2605.28014v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) improves the reasoning performance of large language models (LLMs…
通过pass-rate加权自蒸馏,恢复LLM推理的“甜蜜点”,破解GRPO归一化带来的学习偏差。
arXiv:2605.27765v1 Announce Type: cross Abstract: Self-Distillation Policy Optimization (SDPO) provides dense token-level credit assignment for reinfo…
视频推理新突破:结构化自蒸馏方法VISD可显著提升模型理解能力
arXiv:2605.06094v4 Announce Type: replace-cross Abstract: Training VideoLLMs for complex reasoning remains challenging due to sparse sequence level re…
一种针对大模型推理能力的自适应蒸馏方法,根据模型当前水平动态调整教学方向,提升效果与效率。
arXiv:2605.22263v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) is an emerging LLM post-training paradigm in which the model serv…
无需外部信号,自我蒸馏新范式:能力选择子空间投影让LLM自主提升性能
arXiv:2605.22675v1 Announce Type: new Abstract: Self-distillation bootstraps large language models (LLMs) by training on their own generations. Howeve…
探索统一自蒸馏框架UniSD,为大型语言模型的高效优化与性能提升提供新思路。
arXiv:2605.06597v2 Announce Type: replace Abstract: Self-distillation (SD) offers a promising path for adapting large language models (LLMs) without r…
揭秘自蒸馏为何会损害LLM的数学推理能力,并指出抑制关键探索过程是背后原因。
arXiv:2603.24472v3 Announce Type: replace-cross Abstract: Self-distillation has emerged as an effective post-training paradigm for LLMs, often improvi…
互补自蒸馏如何维护大模型上下文完整性?这项研究提出双模型协作新方案,为LLM安全对齐提供创新思路。
arXiv:2605.20258v1 Announce Type: new Abstract: Contextual Integrity (CI) defines privacy not merely as keeping information hidden, but as governing i…
最新研究:后训练MoE模型通过自蒸馏跳过一半专家,无需从头预训练,显著降低计算量。
arXiv:2605.18643v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) scales language models efficiently through sparse expert activation, and its …