A Trust-region Framework for Moment Estimation
一份用信赖域框架重新审视自适应矩估计的数学研究,适合想深挖优化器收敛原理的读者。
arXiv:2608.04026v1 Announce Type: cross Abstract: In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment…
一份用信赖域框架重新审视自适应矩估计的数学研究,适合想深挖优化器收敛原理的读者。
arXiv:2608.04026v1 Announce Type: cross Abstract: In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment…
揭示大语言模型在线自我改进对齐的收敛性理论,为LLM安全与性能优化提供数学基础,UAI 2026收录。
arXiv:2606.31524v1 Announce Type: cross Abstract: The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel for…
首次从理论层面证明同质深度网络在持续学习中的收敛性质,为灾难性遗忘提供数学解释
arXiv:2606.30559v1 Announce Type: new Abstract: We characterize weakly regularized continual classification in homogeneous models as sequential projec…
AI最终破解了最古老随机梯度下降算法的复杂度难题,数学理论迎来新突破。
arXiv:2606.29593v1 Announce Type: new Abstract: In 1937, Stefan Kaczmarz proposed a simple algorithm for solving systems of linear equations. This alg…
无需投影的线性TD(0)用单一固定步长即可同时获得稳健与快速收敛,理论突破值得关注。
arXiv:2606.24981v1 Announce Type: new Abstract: We study linear TD(0) under Markovian sampling, where data are generated along a single trajectory. We…
探讨控制问题中贝尔曼残差最小化的几何与收敛理论,为强化学习提供新视角
arXiv:2601.18840v4 Announce Type: replace Abstract: Markov decision problems are most commonly solved via dynamic programming. Another approach is Bel…
首次证明带接口约束的工作流学习能保证收敛,理论突破为自动化流程生成奠基。
arXiv:2605.19140v1 Announce Type: new Abstract: We study workflow learning in a setting where specialized agents hand off control through a shared art…
分布式感知器在有限陈旧性与部分参与场景下的收敛性分析,数学严谨,适合ML理论研究者。
arXiv:2601.10705v3 Announce Type: replace Abstract: We study a semi-asynchronous client-server perceptron trained via iterative parameter mixing (IPM-…