1
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
大模型能否自我优化智能体框架?新基准 HarnessOpt-Bench 给出评估方法,直击 Agent 系统关键短板。
arXiv:2608.06301v1 Announce Type: new Abstract: As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the mo…