1
A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions
从期望奖励到GRPO,用第一性原理推导LLM策略优化,揭示结构性扩展的数学本质,适合算法研究者。
arXiv:2606.16733v1 Announce Type: new Abstract: Policy gradient algorithms for language models optimize the same objective $J(\theta) = \mathbb{E}*{\t…