Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations
大模型评估常浪费采样算力,这项研究用贝叶斯最优停止动态决定何时收手,兼顾精度与成本,是做评测必读的思路。
arXiv:2608.14425v1 Announce Type: new Abstract: LLM evaluations often use fixed sampling budgets, testing every item the same number of times even aft…