Learning Generalizable Behaviors for Terminal Agents
用强化学习让大模型在终端环境中学会通用行为,打开日常自动化新可能。
arXiv:2608.22631v1 Announce Type: new Abstract: Terminal agents are a compelling application of large language models (LLMs), with the potential to in…
用强化学习让大模型在终端环境中学会通用行为,打开日常自动化新可能。
arXiv:2608.22631v1 Announce Type: new Abstract: Terminal agents are a compelling application of large language models (LLMs), with the potential to in…
用强化学习让大模型学会「讲道理」,逆合成预测从黑箱猜测走向可解释推理——AI驱动分子设计又进一步。
arXiv:2507.17448v2 Announce Type: replace-cross Abstract: Retrosynthetic planning is a cornerstone of organic synthesis and drug discovery. Yet existi…
多轮推理强化学习新方案,细粒度奖励结构直击信用分配难题,提升LLM代理决策可靠性。
arXiv:2505.11821v3 Announce Type: replace Abstract: Reinforcement Learning (RL) approaches have been wildly used to enhance the reasoning capabilities…
AI tutor 直接从你的资料出题,读书科研的好帮手,告别泛泛而谈的通用问答
I built Learn Leap because I found myself constantly asking ChatGPT questions while reading research papers. I wanted something that already understoo…
一步错步步错?实测三类信用信号在LLM智能体训练中均无法锁定关键步骤,揭示因果归因盲区。
arXiv:2608.19760v1 Announce Type: new Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld)…
离线强化学习遇上LLM推理,未来策略近似方法带来新突破,AI研究者不容错过。
arXiv:2509.19893v3 Announce Type: replace Abstract: Reinforcement learning (RL) has emerged as a key driver of post-training for complex reasoning in …
只用极少GPU就能微调长时程LLM智能体,破解长程推理训练算力困局的新方案。
arXiv:2608.17310v1 Announce Type: new Abstract: Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon a…
大模型结合强化学习,打造精准适配每位学生的AI导师,个性化教育迎来新突破。
arXiv:2608.16907v1 Announce Type: cross Abstract: Generative AI (GenAI) is rapidly reshaping education by unlocking the potential for personalized tut…
突破性智能体强化学习框架,解锁自主决策新范式,AI研究者必读前沿成果。
arXiv:2608.17528v1 Announce Type: new Abstract: Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the …
应对大模型智能体强化学习训练中的故障难题,为长时训练提供高效容错保障,值得关注。
arXiv:2608.14635v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horiz…
给大模型强化学习装上“特权价值函数”,让模型更懂自己表现,训练效率有望大幅提升。
arXiv:2608.16739v1 Announce Type: new Abstract: Reinforcement learning algorithms for Large Language Models (LLMs) are largely distinguished by their …
面向长时程智能体,TRCA用过渡级评分信用分配破解稀疏奖励难题,值得关注。
arXiv:2608.16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, …
用AI定制从入门到进阶的课程,再通过提问检测你的理解,开源可自部署。
I used to learn from books and courses. But I kept realising I hadn't built a real mental model. I always needed an LLM on the side to fill in what th…
PVC塑料三小时变润滑油,绿色化学新突破,废旧管道也能重获新生!
IT之家 8 月 15 日消息,美国弗吉尼亚理工大学的研究团队开发出一种新工艺,可将聚氯乙烯(PVC)塑料转化为聚 α-烯烃,后者是润滑油关键成分。相关研究成果已于 8 月 5 日发表在《自然》上。 该工艺的研发负责人、弗吉尼亚理工大学化学家刘国良(Guoliang “Greg” Liu)表示,团队…
拍照或输入化学题,秒得完整解题步骤,从单位到结果全程可查,刷题对答案两不误。
Article URL: https://chemistryai.chat Comments URL: https://news.ycombinator.com/item?id=49308749 Points: 1 # Comments: 0
揭秘大模型rollout强化学习中的子空间几何,提出GCPO约束方法,助你理解训练背后的数学本质。
arXiv:2608.11674v1 Announce Type: cross Abstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet th…
用强化学习给AI数据中心“降耗”,实测从单卡到集群的LLM训练功率控制,兼顾性能与能效。
arXiv:2608.11226v1 Announce Type: new Abstract: Reinforcement-learning post-training dominates modern language-model development, yet its power behavi…
异步LLM训练如何对抗数据陈旧?这项研究提出陈旧感知的近端策略近似,让PPO在异步场景下更稳更快。
arXiv:2512.06547v4 Announce Type: replace-cross Abstract: Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the h…
从步骤级推理入手,强化大模型自我纠错能力,为提升推理可靠性提供新思路。
arXiv:2608.11573v1 Announce Type: cross Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a f…
自博弈遇上自指导:新方法以自我引导驱动大规模训练,显著提升智能体策略多样性。
arXiv:2604.20209v2 Announce Type: replace Abstract: LLM self-play algorithms are notable in that, in principle, nothing bounds their learning: a Conje…