1
Q-Learning with Fine-Grained Gap-Dependent Regret
Q-learning新理论突破,细粒度依赖遗憾界优化算法性能。
arXiv:2510.06647v2 Announce Type: replace-cross Abstract: We study fine-grained gap-dependent regret bounds for model-free reinforcement learning in e…
Q-learning新理论突破,细粒度依赖遗憾界优化算法性能。
arXiv:2510.06647v2 Announce Type: replace-cross Abstract: We study fine-grained gap-dependent regret bounds for model-free reinforcement learning in e…
提出首个在重尾MDP上同时实现随机与对抗环境最优遗憾的BoBW算法,突破保守局限。
arXiv:2602.01295v3 Announce Type: replace Abstract: We investigate episodic Markov Decision Processes with heavy-tailed losses (HTMDPs). Existing appr…