Generalized Linear Markov Decision Process
提出广义线性马尔可夫决策过程统一框架,为强化学习理论分析提供新视角,值得关注。
arXiv:2506.00818v2 Announce Type: replace-cross Abstract: Offline reinforcement learning for longitudinal studies often faces two linked challenges: r…
提出广义线性马尔可夫决策过程统一框架,为强化学习理论分析提供新视角,值得关注。
arXiv:2506.00818v2 Announce Type: replace-cross Abstract: Offline reinforcement learning for longitudinal studies often faces two linked challenges: r…
提出贝尔曼-泰勒分数解码方法,解决状态依赖可行动作集下MDP的优化难题,为强化学习提供新视角。
arXiv:2606.10979v1 Announce Type: new Abstract: Many Markov decision processes (MDPs) in operations research have feasible actions that are state depe…
从封闭马尔可夫决策过程跳脱,用组合视角打造强化学习的新型收缩反馈语义。
arXiv:2605.24759v1 Announce Type: new Abstract: Discounted reinforcement learning is usually presented through Bellman equations on closed Markov deci…
提出首个在重尾MDP上同时实现随机与对抗环境最优遗憾的BoBW算法,突破保守局限。
arXiv:2602.01295v3 Announce Type: replace Abstract: We investigate episodic Markov Decision Processes with heavy-tailed losses (HTMDPs). Existing appr…