The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works
给LLM智能体喂"预判未来"的奖励,反而把策略喂崩了——这项研究戳破了密集奖励的幻觉,揭示了GRPO训练中真正有效的方法。
arXiv:2607.21273v1 Announce Type: new Abstract: Dense per-step supervision is an appealing remedy for sparse-reward, long-horizon LLM agents: reward t…