1
Batched Bandits with Heavy-Tailed Rewards
当奖励分布出现“重尾”异常值,批量决策的赌博机问题如何逼近最优?全新算法与理论分析一次讲透。
arXiv:2510.03798v3 Announce Type: replace Abstract: The batched multi-armed bandit (MAB) problem, where rewards are collected in batches, is pivotal i…