1
Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
用二元提问替代打分评判,让大模型评估更可解释,还能自我改进,值得研究。
arXiv:2606.27226v1 Announce Type: new Abstract: Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexi…