An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning
从EM视角透视RL训练推理模型,揭示PPO/GRPO的数学本质,为优化大模型推理提供新思路。
arXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabi…