韩国国产自杀式无人机海上试飞失败:两次起飞均数秒后坠海
IT之家 8 月 26 日消息,据新华社今日报道,韩国海军于 8 月 25 日在韩国南部海域进行了国产自杀式无人机(游荡弹药)的海上作战试验,两次放飞均以无人机坠海告终,试验未能完成预定目标。 此次测试在韩国南部海域进行,以大型运输舰“马罗岛”号(ROKS Marado,满载约 1.45 万吨)为平…
IT之家 8 月 26 日消息,据新华社今日报道,韩国海军于 8 月 25 日在韩国南部海域进行了国产自杀式无人机(游荡弹药)的海上作战试验,两次放飞均以无人机坠海告终,试验未能完成预定目标。 此次测试在韩国南部海域进行,以大型运输舰“马罗岛”号(ROKS Marado,满载约 1.45 万吨)为平…
马斯克罕见内部讲话曝光:亲承Grok落后、披露600亿收购Cursor内幕,还预告用员工数据训练AI,野心与争议齐飞。
IT之家 8 月 25 日消息,当地时间 8 月 14 日,SpaceX 宣布完成了对 AI 编程初创公司 Cursor 的 600 亿美元 (IT之家注:现汇率约合 4,043.23 亿元人民币) 收购。 据 The Information 报道,就在官宣当天,埃隆 · 马斯克(Elon Musk…
系统化诊断长时程安全LLM代理的失败根因,突破端到端指标的盲区。
arXiv:2608.20563v1 Announce Type: cross Abstract: Long-horizon security LLM agents must carry information and decisions across many dependent interact…
AI幻觉率仅0.2%,却让生产流程空转245次,一场关于AI可靠性的真实事故复盘。
We run ~100 LLM agents unattended on local models. Last week we found one document that had been rewritten 245 times in 5 days — every attempt rejecte…
面对LLM代理频繁修订判断标准,这项研究揭示失败模式并提出追踪锚定协议,值得AI研究者细读。
arXiv:2608.20729v1 Announce Type: new Abstract: Language-model agents can improve after failure or carry text across episodes without revising what co…
IT之家 8 月 23 日消息,据法治天地今日报道,今年 7 月 8 日,某互联网企业向上海黄浦警方报案,称某自媒体账号发布内容为“知名企业上市失败”的不实贴文,对公司声誉造成负面影响。 黄浦公安分局网安部门接报后迅速开展调查,锁定了该账号运营者莫某。经查,莫某在某证券投资平台运营一个投资类账号,他…
不靠大模型推理,轻量图神经网络也能精准归因智能体失败,效率与可解释性兼得。
arXiv:2608.18575v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems (MAS) often exhibit complex failure modes, which …
给自主智能体戴上“安全紧箍”,行动边界+可信溯源+失败即停,杜绝越权操作。
arXiv:2608.16891v1 Announce Type: new Abstract: Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change w…
IT之家 8 月 19 日消息,蓝箭航天朱雀三号遥二运载火箭 今日在东风商业航天创新试验区发射升空 ,火箭一子级按预定程序成功着陆于甘肃省民勤县朱雀三号着陆场坪,飞行任务取得圆满成功。这是我国在重复使用火箭关键技术上取得的又一重大突破,朱雀三号成为我国首款成功入轨并实现陆地回收的运载火箭。 据央视新…
大模型能预判推理失败,却难选协作协议?这篇论文提出成本感知路由,高效平衡推理性能与开销。
arXiv:2608.14927v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but…
别迷信Agent技能!实证研究发现,某些技能反而会拖累大模型任务成功率,反直觉结论值得关注。
arXiv:2608.11888v1 Announce Type: new Abstract: Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can s…
不用等生成完,从预生成激活就能预判模型成败,给LLM装上“事前质检仪”。
arXiv:2602.09924v4 Announce Type: replace-cross Abstract: Running LLMs with extended reasoning on every problem is expensive, but determining which in…
LLM推理并非万能,主观任务中会失灵?这篇论文剖析失败模式并给出动态路由方案。
arXiv:2608.08889v1 Announce Type: new Abstract: Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a…
从真实翻车中提炼的20条Claude Code技能,专治“看起来会了”的AI编程幻觉。
Article URL: https://whetstone.akbarsha.dev/ Comments URL: https://news.ycombinator.com/item?id=49235255 Points: 3 # Comments: 0
长时程搜索智能体频繁翻车?这个审计框架能精准定位并归因每一步失败。
arXiv:2608.05212v1 Announce Type: new Abstract: Deep search agents tackle challenging questions through long-horizon web interactions, a process that …
多语言环境下多智能体规划为何频频“翻车”?这篇论文给出可操作的失效诊断,直击任务关键信息丢失的根源。
arXiv:2608.03735v1 Announce Type: cross Abstract: Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rare…
从思维链的动态变化中识别大模型推理失误,为AI安全与可解释性提供全新监测路径,值得关注。
arXiv:2608.03291v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing …
大模型意图分类新研究:对比零样本与微调路线,揭示精度、鲁棒性与失败边界,值得关注。
arXiv:2608.02415v1 Announce Type: new Abstract: Intent classification in Large Language Models (LLMs) involves categorizing user prompts into predefin…
用超图建模LLM推理失败模式,成对归因精准定位错误根源,提升可解释性。
arXiv:2608.02026v1 Announce Type: new Abstract: Reflection is a powerful mechanism for LLM reasoning, yet its effectiveness hinges on accurately attri…
为LLM智能体构建自适应失败分类体系,系统提升多智能体场景的可靠性与调试效率。
Article URL: https://multi-agent-systems-failure-taxonomy.github.io/AdaMAST/blogs/adamast_paper/ Comments URL: https://news.ycombinator.com/item?id=49…