Distributionally Robust and Safe Imitation Learning
分布鲁棒性与安全性结合,提升模仿学习在不确定环境下的表现
arXiv:2607.13436v1 Announce Type: new Abstract: Imitation learning (IL) has achieved remarkable success in complex decision-making tasks. However, its…
When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon
一篇挑战LLM后训练中在线模仿学习优势的论文,深入剖析了非可实现性与时间跨度的关键作用
arXiv:2606.30445v1 Announce Type: new Abstract: Online imitation learning (IL), particularly on-policy distillation, has emerged as a strong LLM post-…
Difference-Aware Retrieval Policies for Imitation Learning
一键直达ICLR 2026最新模仿学习论文,包含代码和演示,助你快速掌握差异感知检索策略
arXiv:2606.09758v1 Announce Type: cross Abstract: Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-dis…
Think Like a Pilot: Fine-Grained Long-Horizon UAV Navigation
受飞行员决策启发,提出精细长时域无人机导航方法,提升复杂环境下的自主飞行能力。
arXiv:2606.06836v1 Announce Type: cross Abstract: Language-guided UAV agents must execute long-horizon semantic instructions while producing smooth, p…
Two is better than one: A Collapse-free Multi-Reward RLIF Training Framework
破解多奖励强化学习中的模型崩溃难题,提出RLIF训练新框架实现稳定收敛。
arXiv:2605.22620v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning abili…
Building Deep Graph Predictors with Graph Imitation Learning
快速访问图模仿学习前沿论文,支持PDF在线阅读与下载,助你掌握深度图预测最新方法
arXiv:2601.15133v3 Announce Type: replace-cross Abstract: Recent years have seen substantial progress in neural generation of text, images, and audio,…
IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation
提出短视意图模型IntentVLA,解决机器人模仿数据多模态歧义问题,提升操作准确性
arXiv:2605.14712v1 Announce Type: cross Abstract: Robot imitation data are often multimodal: similar visual-language observations may be followed by d…
An Introduction to Deep Reinforcement and Imitation Learning
系统介绍深度强化学习与模仿学习,面向具身智能体复杂决策问题,是入门该领域的高质量综述。
arXiv:2512.08052v3 Announce Type: replace-cross Abstract: Embodied agents, such as robots and virtual characters, must continuously select actions to …