PATH-Bench: Path-Dependent Evaluation of Lifelong Agents
首次从路径依赖视角评估终身学习智能体,为AI长期能力研究提供新基准。
arXiv:2608.01149v1 Announce Type: new Abstract: Lifelong LLM agents increasingly adapt through external learning states that store past interactions a…
首次从路径依赖视角评估终身学习智能体,为AI长期能力研究提供新基准。
arXiv:2608.01149v1 Announce Type: new Abstract: Lifelong LLM agents increasingly adapt through external learning states that store past interactions a…
零样本评估空中多模态大模型智能体任务能力,无人机AI新突破
arXiv:2607.22014v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, y…
一套评估智能体生物安全风险的双指标框架,兼顾性能与安全对齐,为AI治理提供量化工具。
arXiv:2607.05462v1 Announce Type: cross Abstract: As AI agents are incorporated into life science workflows, the capabilities that speed discovery mig…
首个面向数据智能体的综合基准测试,涵盖多维度评估,推动Agent在数据分析场景落地。
arXiv:2607.01647v1 Announce Type: cross Abstract: Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of …
让AI学会自我评估技能运用,动态进化评分标准,精准提升智能体能力。
arXiv:2607.01874v1 Announce Type: new Abstract: Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool wor…
评估工具使用型AI智能体的可执行基准套件,提供标准化测试场景与量化指标,助力模型优化。
arXiv:2605.11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-…
用千级工具场景拷问LLM智能体规划能力,工具生态越复杂越见真章。
arXiv:2606.22388v1 Announce Type: new Abstract: LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering r…
新基准SkillsBench,系统评估AI Agent技能在不同任务间的通用性与迁移能力,为智能体研究提供重要参考。
arXiv:2602.12670v4 Announce Type: replace Abstract: Agent Skills are structured packages of procedural knowledge that augment large language model (LL…
揭示LLM在调用陌生API时的知识短板,推出首个面向新型API获取能力的智能体评测基准。
arXiv:2606.03657v1 Announce Type: new Abstract: Large language models for code generation often need to use APIs that are absent from their pretrainin…