1
EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts
针对强化学习训练中 rollout 生成缓慢的瓶颈,提出了一种系统感知的自推测解码方法,显著加速 LLM 推理。
arXiv:2606.18967v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a representative post-training paradigm for LLMs, enabling stro…