Recursive Synthesis for Long-Horizon Terminal Tasks
面向长时程终端任务,递归合成方法带来全新解决思路,值得关注。
arXiv:2608.05466v1 Announce Type: new Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hun…
面向长时程终端任务,递归合成方法带来全新解决思路,值得关注。
arXiv:2608.05466v1 Announce Type: new Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hun…
用轨迹生成辅助记忆,多模态推理能力再进一步,值得关注的新方法。
arXiv:2608.01922v1 Announce Type: new Abstract: Multimodal Large Reasoning Models (MLRMs) have achieved strong performance on tasks requiring visual u…
用自蒸馏技术增强基于评分标准的强化学习,提升训练效果和泛化能力。
arXiv:2607.18082v1 Announce Type: cross Abstract: Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognize…
先想象再预测:论文提出交错潜在视觉推理新方法,提升视频事件预测准确性与可解释性。
arXiv:2606.05769v1 Announce Type: new Abstract: Video event prediction (VEP) requires models to infer unobserved future states from partial video evid…