1
From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training
提出利用事后经验将结果转化为行动,显著提升长周期语言智能体训练效果,ICML 2026录用论文。
arXiv:2607.16257v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a widely adopted technique for improving large language model…