GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs
揭秘大模型rollout强化学习中的子空间几何,提出GCPO约束方法,助你理解训练背后的数学本质。
arXiv:2608.11674v1 Announce Type: cross Abstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet th…