1
Learning the Supports for Categorical Critic in Reinforcement Learning
强化学习前沿:RLC 2026论文提出分类评论器支撑集学习方法,直击连续控制任务中价值估计不稳定的痛点。
arXiv:2607.01880v1 Announce Type: new Abstract: Value functions are an essential component in actor-critic based deep reinforcement learning (RL). Con…