消息称《GTA 6》大幅强化偷车系统:部分载具游戏初期无法窃取、汽车配有追踪器
IT之家 8 月 29 日消息,R 星开发负责人 Rob Nelson 在接受 IGN 采访时透露,《GTA 6》大幅强化偷车系统,过去玩家只需要按下按键即可轻松开走汽车,但新作中部分载具在游戏初期无法直接盗取,玩家需要随着剧情推进获得相应工具后才能顺利劫车。 Rob Nelson 表示,在游戏之初…
IT之家 8 月 29 日消息,R 星开发负责人 Rob Nelson 在接受 IGN 采访时透露,《GTA 6》大幅强化偷车系统,过去玩家只需要按下按键即可轻松开走汽车,但新作中部分载具在游戏初期无法直接盗取,玩家需要随着剧情推进获得相应工具后才能顺利劫车。 Rob Nelson 表示,在游戏之初…
IT之家 8 月 27 日消息,在接受《Dazed》采访时, Rockstar North 开发负责人罗布 · 纳尔逊(Rob Nelson)透露《侠盗猎车手 VI》(GTA 6)游戏机制,并谈到了两名游戏主角的角色设计。 纳尔逊透露杰森(Jason)和露西娅(Lucia)两位主角的身体会根据玩家选…
用强化学习让大模型在终端环境中学会通用行为,打开日常自动化新可能。
arXiv:2608.22631v1 Announce Type: new Abstract: Terminal agents are a compelling application of large language models (LLMs), with the potential to in…
用强化学习让大模型学会「讲道理」,逆合成预测从黑箱猜测走向可解释推理——AI驱动分子设计又进一步。
arXiv:2507.17448v2 Announce Type: replace-cross Abstract: Retrosynthetic planning is a cornerstone of organic synthesis and drug discovery. Yet existi…
多轮推理强化学习新方案,细粒度奖励结构直击信用分配难题,提升LLM代理决策可靠性。
arXiv:2505.11821v3 Announce Type: replace Abstract: Reinforcement Learning (RL) approaches have been wildly used to enhance the reasoning capabilities…
OpenAI改口支持加州AI安全法案,监管态度大转弯背后有何考量?
IT之家 8 月 23 日消息,OpenAI 呼吁加州进一步加强一项具有里程碑意义的 AI 安全法案。该法案已于去年获得通过。 OpenAI 全球事务团队在 LinkedIn 上发表的一篇文章中表示,加州 SB 53 法案“应该进一步扩大安全保障措施”。例如,可以要求对正在训练或评估的前沿 AI 模…
一步错步步错?实测三类信用信号在LLM智能体训练中均无法锁定关键步骤,揭示因果归因盲区。
arXiv:2608.19760v1 Announce Type: new Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld)…
离线强化学习遇上LLM推理,未来策略近似方法带来新突破,AI研究者不容错过。
arXiv:2509.19893v3 Announce Type: replace Abstract: Reinforcement learning (RL) has emerged as a key driver of post-training for complex reasoning in …
只用极少GPU就能微调长时程LLM智能体,破解长程推理训练算力困局的新方案。
arXiv:2608.17310v1 Announce Type: new Abstract: Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon a…
视频多模态模型常“看走眼”?这项研究用结构化奖励强化推理一致性,为视频理解立新标尺。
arXiv:2604.01460v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding.…
大模型结合强化学习,打造精准适配每位学生的AI导师,个性化教育迎来新突破。
arXiv:2608.16907v1 Announce Type: cross Abstract: Generative AI (GenAI) is rapidly reshaping education by unlocking the potential for personalized tut…
突破性智能体强化学习框架,解锁自主决策新范式,AI研究者必读前沿成果。
arXiv:2608.17528v1 Announce Type: new Abstract: Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the …
应对大模型智能体强化学习训练中的故障难题,为长时训练提供高效容错保障,值得关注。
arXiv:2608.14635v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horiz…
给大模型强化学习装上“特权价值函数”,让模型更懂自己表现,训练效率有望大幅提升。
arXiv:2608.16739v1 Announce Type: new Abstract: Reinforcement learning algorithms for Large Language Models (LLMs) are largely distinguished by their …
面向长时程智能体,TRCA用过渡级评分信用分配破解稀疏奖励难题,值得关注。
arXiv:2608.16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, …
集中管理Word、PPT、Excel等资料,AI精准分析摘要,支持Markdown让笔记更灵活。
IT之家 8 月 15 日消息,科技媒体 Neowin 昨日(8 月 14 日)发布博文,报道称微软更新 Copilot Notebooks 应用, 新增支持 Markdown、纯文本和富文本 3 种文本格式。 IT之家注:Copilot Notebooks 是微软 Microsoft 365 Co…
揭秘大模型rollout强化学习中的子空间几何,提出GCPO约束方法,助你理解训练背后的数学本质。
arXiv:2608.11674v1 Announce Type: cross Abstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet th…
用强化学习给AI数据中心“降耗”,实测从单卡到集群的LLM训练功率控制,兼顾性能与能效。
arXiv:2608.11226v1 Announce Type: new Abstract: Reinforcement-learning post-training dominates modern language-model development, yet its power behavi…
异步LLM训练如何对抗数据陈旧?这项研究提出陈旧感知的近端策略近似,让PPO在异步场景下更稳更快。
arXiv:2512.06547v4 Announce Type: replace-cross Abstract: Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the h…
从步骤级推理入手,强化大模型自我纠错能力,为提升推理可靠性提供新思路。
arXiv:2608.11573v1 Announce Type: cross Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a f…