1
Robust and Fast Training via Per-Sample Clipping
逐样本裁剪的稳健梯度估计,同时提速训练并增强鲁棒性,值得一试。
arXiv:2605.02701v2 Announce Type: replace-cross Abstract: We propose a robust gradient estimator based on per-sample gradient clipping and analyze its…
逐样本裁剪的稳健梯度估计,同时提速训练并增强鲁棒性,值得一试。
arXiv:2605.02701v2 Announce Type: replace-cross Abstract: We propose a robust gradient estimator based on per-sample gradient clipping and analyze its…
被ICML接收的自适应梯度裁剪方法,有效提升LLM预训练稳定性,AI训练优化的新突破
arXiv:2502.11034v3 Announce Type: replace Abstract: Loss spikes remain a persistent obstacle in large-scale language model pretraining. While previous…
MuCon通过裁剪奇异值改进Muon优化器,为LLM训练提供更稳定的更新策略
arXiv:2605.26459v1 Announce Type: new Abstract: Muon-style optimizers take a matrix-valued momentum or preconditioned update $B = U \operatorname{diag…