Training Proactive and Personalized LLM Agents
让大模型从“被动应答”走向“主动服务”,教你训练会察言观色的个性化AI代理,附完整方法框架。
arXiv:2511.02208v2 Announce Type: replace Abstract: Despite rapid progress, current AI agents are primarily optimized for isolated task completion. We…
让大模型从“被动应答”走向“主动服务”,教你训练会察言观色的个性化AI代理,附完整方法框架。
arXiv:2511.02208v2 Announce Type: replace Abstract: Despite rapid progress, current AI agents are primarily optimized for isolated task completion. We…
揭秘大模型智能体如何通过工具实现“遗忘”,为安全与能力平衡提供新思路。
arXiv:2608.21544v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can…
不再只关注序列预测,而是引入关系不确定性传播,让LLM智能体在结构化决策中更可靠,值得深入研读。
arXiv:2608.16002v1 Announce Type: cross Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agent…
把软件交付全流程塞进GitHub,Agent应用让自动化一步到位
See how four GitHub Agent Apps can help you scope, secure, roll out, and ship a feature across the SDLC–all without leaving GitHub. The post How to br…
新工具OliverGraph连接Slack、GitHub和文档,为AI代理补充工作上下文,让智能体更懂真实业务。
We just launched OliverGraph in beta. OliverGraph connects to tools like Slack, GitHub, and docs and keeps track of the context behind the work happen…
当AI代理多到找不着,这家“工厂”专治选择困难,精准检索领域专属智能体。
arXiv:2608.09934v1 Announce Type: new Abstract: Large language model (LLM) agents improve task performance by decomposing problems into role-specializ…
给真正交付的团队:AGENTS.md不是README,而是代码库地图。用5-10条要点、300行上限,让AI代理精准理解你的业务逻辑。
By the end of 2025, AI coding agents stopped being a toy. Claude Code, Codex, Cursor, Copilot and opencode are now doing real work in real repos — and…
LLM智能体自我进化可能反噬,预承诺门控机制防止技能污染,揭秘新解法。
arXiv:2608.05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajecto…
为LLM智能体打造风险感知世界模型的运行时护栏,用世界模型预判危险,安全高效两不误。
arXiv:2608.05695v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly invoke external tools and interact with real-world s…
流式任务下自进化智能体表现如何?新基准AgentStream揭示性能边界与进化策略。
arXiv:2608.00155v1 Announce Type: new Abstract: Large language model (LLM) agents can self-evolve by continually improving from their own accumulated …
长程任务中LLM智能体记忆混乱?这项决策时记忆仲裁机制让智能体精准调用关键信息。
arXiv:2608.02113v1 Announce Type: new Abstract: Large language model (LLM) agents must retain and use cross-step information to act coherently in long…
别再只调提示词了,学会用上下文工程驱动AI Agent,让交互更聪明高效。
AI agents get talked about a lot, but most explanations stay abstract. Here's a short, practical breakdown of what actually makes an agent work — plus…
AI 智能体为何需要“技能”?用 Markdown 文档真的够吗?HN 开发者吵出了架构真相。
I’ve asked AI a few times and I still don’t get it. Why do frameworks like Claude Code or Codex have the concept of “skills” instead of just using wel…
用WebCMD范式打造毫秒级Hacker News CLI,让AI代理绕过浏览器开销,延迟降95%,token省90%。
How we leveraged the WebCMD paradigm to slash web-navigation latency by 95%, reduce LLM token consumption by 90%, and establish a deterministic, schem…
为LLM智能体构建自适应失败分类体系,系统提升多智能体场景的可靠性与调试效率。
Article URL: https://multi-agent-systems-failure-taxonomy.github.io/AdaMAST/blogs/adamast_paper/ Comments URL: https://news.ycombinator.com/item?id=49…
开源快照测试工具,记录AI Agent的LLM和工具调用,支持语义回归测试。
Article URL: https://github.com/iamfaham/AgentSnap Comments URL: https://news.ycombinator.com/item?id=49099464 Points: 4 # Comments: 0
提出分层技能图框架,让LLM代理通过结构化技能更高效协作与决策。
arXiv:2607.25853v1 Announce Type: new Abstract: Skills have become an important abstraction for enabling large language model (LLM) agents to reuse pa…
AI agent如何颠覆传统分层架构?作者分享转向垂直切分的新认知。
The old preference For years, I preferred a traditional layered architecture. Controllers/ Services/ Repositories/ DTOs/ Models/ It makes navigation s…
重磅论文揭露:主流LLM Agent评估框架忽略静默故障,提出轻量黑盒审计方案,精准检测恶意拒绝漏洞
arXiv:2607.19449v1 Announce Type: new Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or expl…
开源macOS终端Yorishiro,让AI agents以3D角色形式活跃在终端里,体验前所未有的交互方式。
I’ve been building Yorishiro, an open-source macOS terminal for working with coding agents like Claude Code and Codex. The name is Japanese — a yorish…