(How) Learning Rates Regulate Catastrophic Overtraining
学习率如何抑制灾难性过训练?这项研究揭示关键调节机制,为深度学习训练提供全新视角。
arXiv:2604.13627v2 Announce Type: replace Abstract: Supervised fine-tuning (SFT) is a common first stage of LLM post-training, teaching the model to f…
学习率如何抑制灾难性过训练?这项研究揭示关键调节机制,为深度学习训练提供全新视角。
arXiv:2604.13627v2 Announce Type: replace Abstract: Supervised fine-tuning (SFT) is a common first stage of LLM post-training, teaching the model to f…
揭秘LLM训练中学习率缩放的非线性规律,为优化大模型训练过程提供全新理论视角。
arXiv:2606.29158v1 Announce Type: new Abstract: Learning-rate transfer can reduce the cost of training large language models: instead of sweeping lear…
LLM继续预训练中,超参数配置可预测的缩放规律,告别启发式搜索与高昂成本
arXiv:2606.05610v1 Announce Type: new Abstract: The efficacy of continued pre-training for Large Language Models (LLMs) hinges upon hyperparameter con…
打破传统训练调度限制,谱优化实现随时中断与继续,灵活高效。
arXiv:2605.23061v1 Announce Type: cross Abstract: Standard neural network training relies on learning-rate schedules tied to a fixed horizon, leading …
突破传统统一学习率,重尾分布指导LLM逐层自适应学习,大幅提升训练效率与模型性能。
arXiv:2605.22297v1 Announce Type: cross Abstract: Learning rate configuration is a fundamental aspect of modern deep learning. The prevailing practice…
揭秘SGD在LLM预训练中不如Adam的根源:大有效学习率的关键作用。
arXiv:2605.17787v1 Announce Type: new Abstract: It is widely believed that stochastic gradient descent (SGD) performs significantly worse than adaptiv…
通过调整学习率,简单LoRA即可媲美复杂微调方法,揭示被忽视的关键因素。
arXiv:2602.04998v2 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) is the prevailing approach for efficient large language model (LLM) fin…