1
Understanding Knowledge Distillation in Post-Training: When It Helps and When It Fails
不只看蒸馏提效,更揭示失效边界,给后训练选型提供实操判断依据
arXiv:2606.22942v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance across many tasks, but their high computationa…