1
Repeated post-training is not Self-improving: Diagnosing Scientific Amnesia in Continual DPO Pipelines
反复后训练并非自我提升?新研究揭示DPO持续训练中的“科学遗忘”现象,值得关注。
arXiv:2606.21089v1 Announce Type: new Abstract: Industrial LLM teams often ship behavior updates by repeatedly DPO-training a base model on sequences …