Batched Bandits with Heavy-Tailed Rewards
当奖励分布出现“重尾”异常值,批量决策的赌博机问题如何逼近最优?全新算法与理论分析一次讲透。
arXiv:2510.03798v3 Announce Type: replace Abstract: The batched multi-armed bandit (MAB) problem, where rewards are collected in batches, is pivotal i…
当奖励分布出现“重尾”异常值,批量决策的赌博机问题如何逼近最优?全新算法与理论分析一次讲透。
arXiv:2510.03798v3 Announce Type: replace Abstract: The batched multi-armed bandit (MAB) problem, where rewards are collected in batches, is pivotal i…
把抽样设计的无偏性引入算法机器学习,为统计推断与模型训练架起新桥梁,值得技术党细读。
arXiv:2606.28795v1 Announce Type: new Abstract: Machine Learning (ML) algorithms, such as k-Nearest Neighbours (kNN) or random forest, eschew the idea…
探索在线战略分类中随机化算法的新进展,揭示其如何应对博弈分类场景
arXiv:2602.06257v2 Announce Type: replace Abstract: Online strategic classification studies settings in which agents strategically modify their featur…
均值算法理论突破,揭示其下界与遗憾的互补关系。
arXiv:2606.04931v1 Announce Type: new Abstract: Mean-based algorithms are a class of online learning algorithms that assign low probability to actions…
新论文提出通过logits凸性稳定策略优化,为强化学习训练提供理论新视角。
arXiv:2603.00963v2 Announce Type: replace Abstract: While reinforcement learning (RL) has been central to the recent success of large language models …
首次为全局优化问题提供理论保证的Proximal basin hopping算法,打破传统纯启发式方法局限
arXiv:2605.18364v1 Announce Type: new Abstract: Global optimization is a challenging problem, with plenty of algorithms displaying empirical success, …