Benchmarking LLM Judges for Mobile Agent Evaluation
移动代理评估中LLM裁判靠谱吗?全新基准MobileJudgeBench首次系统检验其可靠性。
arXiv:2608.11434v1 Announce Type: new Abstract: Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the rel…