1
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
大模型当裁判时总护着自己?这项研究用激活干预破解自偏好,让AI评分更公平。
arXiv:2509.03647v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer fro…