Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
用离线Top-K logits和融合分块KL损失,大幅降低大模型蒸馏成本,兼顾效率与质量。
arXiv:2608.03796v1 Announce Type: cross Abstract: Small language models are often the only option for deployment under tight latency, cost, and on-pre…