Learning Generalizable Behaviors for Terminal Agents
用强化学习让大模型在终端环境中学会通用行为,打开日常自动化新可能。
arXiv:2608.22631v1 Announce Type: new Abstract: Terminal agents are a compelling application of large language models (LLMs), with the potential to in…
用强化学习让大模型在终端环境中学会通用行为,打开日常自动化新可能。
arXiv:2608.22631v1 Announce Type: new Abstract: Terminal agents are a compelling application of large language models (LLMs), with the potential to in…
用强化学习让大模型学会「讲道理」,逆合成预测从黑箱猜测走向可解释推理——AI驱动分子设计又进一步。
arXiv:2507.17448v2 Announce Type: replace-cross Abstract: Retrosynthetic planning is a cornerstone of organic synthesis and drug discovery. Yet existi…
多轮推理强化学习新方案,细粒度奖励结构直击信用分配难题,提升LLM代理决策可靠性。
arXiv:2505.11821v3 Announce Type: replace Abstract: Reinforcement Learning (RL) approaches have been wildly used to enhance the reasoning capabilities…
一步错步步错?实测三类信用信号在LLM智能体训练中均无法锁定关键步骤,揭示因果归因盲区。
arXiv:2608.19760v1 Announce Type: new Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld)…
离线强化学习遇上LLM推理,未来策略近似方法带来新突破,AI研究者不容错过。
arXiv:2509.19893v3 Announce Type: replace Abstract: Reinforcement learning (RL) has emerged as a key driver of post-training for complex reasoning in …
只用极少GPU就能微调长时程LLM智能体,破解长程推理训练算力困局的新方案。
arXiv:2608.17310v1 Announce Type: new Abstract: Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon a…
大模型结合强化学习,打造精准适配每位学生的AI导师,个性化教育迎来新突破。
arXiv:2608.16907v1 Announce Type: cross Abstract: Generative AI (GenAI) is rapidly reshaping education by unlocking the potential for personalized tut…
突破性智能体强化学习框架,解锁自主决策新范式,AI研究者必读前沿成果。
arXiv:2608.17528v1 Announce Type: new Abstract: Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the …
应对大模型智能体强化学习训练中的故障难题,为长时训练提供高效容错保障,值得关注。
arXiv:2608.14635v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horiz…
给大模型强化学习装上“特权价值函数”,让模型更懂自己表现,训练效率有望大幅提升。
arXiv:2608.16739v1 Announce Type: new Abstract: Reinforcement learning algorithms for Large Language Models (LLMs) are largely distinguished by their …
面向长时程智能体,TRCA用过渡级评分信用分配破解稀疏奖励难题,值得关注。
arXiv:2608.16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, …
揭秘大模型rollout强化学习中的子空间几何,提出GCPO约束方法,助你理解训练背后的数学本质。
arXiv:2608.11674v1 Announce Type: cross Abstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet th…
用强化学习给AI数据中心“降耗”,实测从单卡到集群的LLM训练功率控制,兼顾性能与能效。
arXiv:2608.11226v1 Announce Type: new Abstract: Reinforcement-learning post-training dominates modern language-model development, yet its power behavi…
异步LLM训练如何对抗数据陈旧?这项研究提出陈旧感知的近端策略近似,让PPO在异步场景下更稳更快。
arXiv:2512.06547v4 Announce Type: replace-cross Abstract: Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the h…
从步骤级推理入手,强化大模型自我纠错能力,为提升推理可靠性提供新思路。
arXiv:2608.11573v1 Announce Type: cross Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a f…
自博弈遇上自指导:新方法以自我引导驱动大规模训练,显著提升智能体策略多样性。
arXiv:2604.20209v2 Announce Type: replace Abstract: LLM self-play algorithms are notable in that, in principle, nothing bounds their learning: a Conje…
强化学习训练中混合rollout调度的新突破,跳出前缀局部性束缚,为分布式RL提速提供新思路。
arXiv:2608.11152v1 Announce Type: cross Abstract: Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasi…
提出广义线性马尔可夫决策过程统一框架,为强化学习理论分析提供新视角,值得关注。
arXiv:2506.00818v2 Announce Type: replace-cross Abstract: Offline reinforcement learning for longitudinal studies often faces two linked challenges: r…
用MoE代理模型低成本复现LLM强化学习后训练故障,大幅降低排查成本,直击调试痛点。
arXiv:2608.10823v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training of large language models (LLMs) is computationally intensive…
成立两个月拿下11亿美元融资,River AI想重构强化学习底层堆栈,让企业15分钟跑完复杂训练并省下数倍成本。
River AI, a startup founded by xAI co-founder Igor Babuschkin, has a fascinating vision for personal agents and secured $1.1 billion out of the gate.