1
Robust and Fast Training via Per-Sample Clipping
逐样本裁剪的稳健梯度估计,同时提速训练并增强鲁棒性,值得一试。
arXiv:2605.02701v2 Announce Type: replace-cross Abstract: We propose a robust gradient estimator based on per-sample gradient clipping and analyze its…
逐样本裁剪的稳健梯度估计,同时提速训练并增强鲁棒性,值得一试。
arXiv:2605.02701v2 Announce Type: replace-cross Abstract: We propose a robust gradient estimator based on per-sample gradient clipping and analyze its…
AutoML管道如何挑选损失函数与优化器最佳搭档?这项研究为NNGPT系统找到了稳定训练的配对规律
arXiv:2606.20933v1 Announce Type: new Abstract: The choice of loss function and optimizer is an important decision, that shapes further model training…
揭秘冻结视觉模型训练中噪声数据的陷阱:小损失策略为何失效?跨数据集基准带来新洞察。
arXiv:2605.22591v1 Announce Type: new Abstract: Frozen Vision Foundation Models (VFMs) with lightweight classification heads are increasingly used in …