No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
LLM当裁判不靠谱?这项研究揭示无人类标注时AI评分的致命盲区,做评估必看。
arXiv:2503.05061v3 Announce Type: replace Abstract: Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expand…