1
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning
不止优化均值,多矩策略让LLM推理训练更稳健,值得一读的强化学习新思路。
arXiv:2608.02149v1 Announce Type: new Abstract: Reinforcement learning has become a central paradigm for improving the reasoning capabilities of large…