Predictable GRPO: A Closed-Form Model of Training Dynamics
首次用闭式模型刻画GRPO训练动态,为强化学习调参提供可预测的理论框架。
arXiv:2606.30789v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard tool for improving the reasoning abili…
首次用闭式模型刻画GRPO训练动态,为强化学习调参提供可预测的理论框架。
arXiv:2606.30789v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard tool for improving the reasoning abili…
揭秘Transformer训练中的动态机制,为AI可解释性提供新视角,大模型研究者必读。
arXiv:2410.24050v3 Announce Type: replace Abstract: Large-scale pretraining of transformers has been central to the success of foundation models. Howe…
用标量捕捉神经网络训练动态,简化复杂训练过程的可视化与诊断。
arXiv:2606.30384v1 Announce Type: new Abstract: Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimension…
深度揭秘权重衰减对训练稳定性的真实作用,挑战传统正则化认知。
arXiv:2605.16622v1 Announce Type: new Abstract: In modern deep learning, weight decay is often credited with "stabilizing" training dynamics, divergin…