1
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?
首个评估LLM Agent端到端解决运筹学任务的基准,揭示大模型在复杂优化问题上的能力边界
arXiv:2606.19787v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents for multi-step tasks in executabl…