Training Proactive and Personalized LLM Agents
让大模型从“被动应答”走向“主动服务”,教你训练会察言观色的个性化AI代理,附完整方法框架。
arXiv:2511.02208v2 Announce Type: replace Abstract: Despite rapid progress, current AI agents are primarily optimized for isolated task completion. We…
让大模型从“被动应答”走向“主动服务”,教你训练会察言观色的个性化AI代理,附完整方法框架。
arXiv:2511.02208v2 Announce Type: replace Abstract: Despite rapid progress, current AI agents are primarily optimized for isolated task completion. We…
揭秘大模型智能体如何通过工具实现“遗忘”,为安全与能力平衡提供新思路。
arXiv:2608.21544v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can…
不再只关注序列预测,而是引入关系不确定性传播,让LLM智能体在结构化决策中更可靠,值得深入研读。
arXiv:2608.16002v1 Announce Type: cross Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agent…
当AI代理多到找不着,这家“工厂”专治选择困难,精准检索领域专属智能体。
arXiv:2608.09934v1 Announce Type: new Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specializ…
LLM智能体自我进化可能反噬,预承诺门控机制防止技能污染,揭秘新解法。
arXiv:2608.05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajecto…
为LLM智能体打造风险感知世界模型的运行时护栏,用世界模型预判危险,安全高效两不误。
arXiv:2608.05695v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly invoke external tools and interact with real-world s…
长程任务中LLM智能体记忆混乱?这项决策时记忆仲裁机制让智能体精准调用关键信息。
arXiv:2608.02113v1 Announce Type: new Abstract: Large language model (LLM) agents must retain and use cross-step information to act coherently in long…
为LLM智能体构建自适应失败分类体系,系统提升多智能体场景的可靠性与调试效率。
Article URL: https://multi-agent-systems-failure-taxonomy.github.io/AdaMAST/blogs/adamast_paper/ Comments URL: https://news.ycombinator.com/item?id=49…
提出分层技能图框架,让LLM代理通过结构化技能更高效协作与决策。
arXiv:2607.25853v1 Announce Type: new Abstract: Skills have become an important abstraction for enabling large language model (LLM) agents to reuse pa…
重磅论文揭露:主流LLM Agent评估框架忽略静默故障,提出轻量黑盒审计方案,精准检测恶意拒绝漏洞
arXiv:2607.19449v1 Announce Type: new Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or expl…
用大模型智能体搞定多人旅行规划,一篇探索LLM Agent在群体决策中应用的论文。
arXiv:2607.18806v1 Announce Type: new Abstract: This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Langu…
揭秘看似无害的工具如何被组装成攻击链,威胁LLM Agent安全,一篇值得关注的漏洞挖掘研究。
arXiv:2509.25624v3 Announce Type: replace-cross Abstract: As LLMs advance into autonomous agents with tool-use capabilities, they introduce security c…
提出公平oracle量化LLM agent“知道却做不到”的认知鸿沟,揭示感知与行动差距根源,值得关注长期决策评估的研究者深读。
arXiv:2607.13618v1 Announce Type: new Abstract: LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost…
两大移动端自动化工具对比:Mobilerun 基于 LLM agent 读懂屏幕并自然语言操作,适合复杂多步任务
Both Mobilerun and DuoPlus put mobile devices in the cloud and let you automate them and both can run many social media accounts at scale with automat…
AI代理自动生成3D游戏可执行的叙事场景,让非专业人士也能创作动态多角色视频。
arXiv:2604.10383v2 Announce Type: replace Abstract: We use LLM agents to author executable specifications for a living world: formal Graphs of Events …
为LLM Agents设计的工作台,支持重现、干预与缓解,提升MCP环境下Agent可靠性。
arXiv:2607.11098v1 Announce Type: cross Abstract: Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a…
为LLM智能体设计可审计的假设演化协议,让AI科学家过程透明可信
arXiv:2607.09195v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly expected to play a central role in AI-driven scient…
用版本号锁定邮件动作,让LLM agent不再自由发挥,提升可预测性与可靠性。
Muchos equipos que integran LLMs con correo se obsesionan con el prompt y dejan medio borroso el contrato de ejecución. En mi experiencia, el fallo re…
ICML 2026论文为LLM Agent在社交困境中的合作能力打造首个系统化基准,揭示维持机制的关键设计。
arXiv:2604.15267v2 Announce Type: replace-cross Abstract: It is increasingly important that LLM agents interact effectively and safely with other goal…
自推测分支技术让LLM agent在等待工具返回时预生成后续推理,大幅减少GPU空闲时间,提升推理效率。
arXiv:2607.03333v1 Announce Type: cross Abstract: LLM agents are becoming a common interface for research, coding, and question answering, yet their T…