1
How Benchmark Prediction from Fewer Data Misses the Mark
用少量数据预测大模型基准表现?这篇论文揭示了偏差与误判背后的关键漏洞。
arXiv:2506.07673v2 Announce Type: replace Abstract: Large language model (LLM) evaluation is increasingly costly, prompting interest in methods that s…