A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
异步LLM训练如何对抗数据陈旧?这项研究提出陈旧感知的近端策略近似,让PPO在异步场景下更稳更快。
arXiv:2512.06547v4 Announce Type: replace-cross Abstract: Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the h…