Agent Lightning v1.0: Towards Harnessed Agentic RL
突破性智能体强化学习框架,解锁自主决策新范式,AI研究者必读前沿成果。
arXiv:2608.17528v1 Announce Type: new Abstract: Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the …
突破性智能体强化学习框架,解锁自主决策新范式,AI研究者必读前沿成果。
arXiv:2608.17528v1 Announce Type: new Abstract: Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the …
提出政策感知训练支架,革新智能体强化学习范式
arXiv:2607.21419v1 Announce Type: new Abstract: In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, produci…
智能体强化学习新框架OpenTinker,通过分离关注点简化训练流程,值得技术研究者细读。
arXiv:2601.07376v2 Announce Type: replace Abstract: We introduce \textsc{OpenTinker}, an open infrastructure for training large language model (LLM) a…
一种自我纯净轨迹方法,大幅提升4B-7B小模型的Agentic强化学习效果,专治执行失败噪声!
arXiv:2601.15141v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (RL) has empowered Large Language Models (LLMs) to utilize tools li…
提出RollArt解耦架构,实现大规模多任务Agentic RL训练,显著提升效率与扩展性。
arXiv:2512.22560v2 Announce Type: replace-cross Abstract: Agentic Reinforcement Learning (RL) trains LLMs through multi-turn interactions with environ…
高效管理Agentic RL后训练资源的新方案Libra,降低训练成本、提升性能。
arXiv:2606.03077v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a standard post-training paradigm for large language models (…