Scheduling Mixed RL Rollouts Beyond Prefix Locality
强化学习训练中混合rollout调度的新突破,跳出前缀局部性束缚,为分布式RL提速提供新思路。
arXiv:2608.11152v1 Announce Type: cross Abstract: Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasi…