1
An Executable Benchmarking Suite for Tool-Using Agents
评估工具使用型AI智能体的可执行基准套件,提供标准化测试场景与量化指标,助力模型优化。
arXiv:2605.11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-…