1
Ask the Right Comparison:Bias-Aware Bayesian Active Top-$k$ Ranking with LLM Judges
LLM裁判偏爱长答案,传统偏置校正为何失效?贝叶斯主动Top-K排序新思路,直击评估偏差核心。
arXiv:2607.02104v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate ou…