End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
强化学习理论新突破:线性Bellman完备MDP在确定性转移下实现端到端高效求解,值得算法研究者细读。
arXiv:2603.23461v2 Announce Type: replace Abstract: We study reinforcement learning (RL) with linear function approximation in Markov Decision Process…