AI with Authority, from Application to Silicon
生成式AI让机器验证从奢侈品变为必需品,成为不可收买的裁判
arXiv:2608.21356v1 Announce Type: cross Abstract: For sixty years, machine verification has been a major cost overhead, affordable only for exceptiona…
生成式AI让机器验证从奢侈品变为必需品,成为不可收买的裁判
arXiv:2608.21356v1 Announce Type: cross Abstract: For sixty years, machine verification has been a major cost overhead, affordable only for exceptiona…
医学影像AI落地最怕模型悄悄“变笨”,MMC+框架用可扩展监控提前捕捉漂移,保障长期可靠性。
arXiv:2410.13174v3 Announce Type: replace-cross Abstract: The integration of artificial intelligence (AI) into medical imaging has advanced clinical d…
前沿大模型存在“响应漂移”?这篇论文系统研究了不同版本LLM输出随时间一致性的关键问题。
arXiv:2607.20454v1 Announce Type: cross Abstract: All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate f…
LLM事实核查的二元判断隐藏可靠性危机,论文提出基于证据链评估的校准选择性方法,直击AI可信度痛点。
arXiv:2607.18240v1 Announce Type: new Abstract: Large language models (LLMs) can achieve strong fact-checking accuracy, yet forced binary decisions co…
用少量数据预测大模型基准表现?这篇论文揭示了偏差与误判背后的关键漏洞。
arXiv:2506.07673v2 Announce Type: replace Abstract: Large language model (LLM) evaluation is increasingly costly, prompting interest in methods that s…
四川俚语被Grok误译成性暗示,一次真实翻车暴露AI跨文化理解短板,专业使用需警惕。
My coworker responded to an invitation to meet up with Sichuan Chinese Venture Capitalists (the event mandates that dialect only), and she used Sichua…
多智能体系统可靠性堪忧?用这个开源工具在投产前揪出协作漏洞
Article URL: https://github.com/surajkumar811/swarm-test Comments URL: https://news.ycombinator.com/item?id=48668678 Points: 1 # Comments: 0
揭示AI模型在静默数据损坏下的脆弱性,为可靠性部署提供关键参考。
arXiv:2405.01741v4 Announce Type: replace-cross Abstract: Reliability of AI systems is a fundamental concern for the successful deployment and widespr…
贝叶斯控制理论为代码生成代理注入概率推理,提升行为可靠性与稳定性
arXiv:2606.24453v1 Announce Type: new Abstract: Modern coding agents pair LLM generators with various tools, including cheap diagnostics and expensive…
AI看似强大却暗藏隐患,一文讲透模型不可靠的根源与应对思路。
Article URL: https://arachnemag.substack.com/p/ais-reliability-gap Comments URL: https://news.ycombinator.com/item?id=48650054 Points: 1 # Comments: 0
科技评论家Cringely联手创办2Brains Inc,专攻大模型“幻觉”难题
Article URL: https://slashdot.org/story/26/06/20/0556251/tech-pundit-cringely-co-founds-startup-2brains-inc-to-solve-llm-hallucinations Comments URL: …
揭示LLM在临床表格数据上的认知盲点,提出跨模型归因分歧检测新方法,提升模型可信度与安全应用。
arXiv:2606.19509v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied to structured clinical data, yet whether they ca…
自主AI系统行为不可预测,传统监控失效,本文揭示了如何通过可观测性深入洞察其内部运作。
Monitoring and Observability for Autonomous AI Systems Autonomous AI systems—from self-driving cars to algorithmic trading bots and robotic process au…
初创公司Probably获投900万美元,打造更可靠的AI错误检测引擎,可扩展至会计医疗等领域。
Probably wants to prevent hallucinations and factual errors from reaching users, and achieve accuracy on par with deterministic systems.
大模型能否胜任医生角色的关键考验:最新研究实探LLM在医疗诊断与临床推理评分中的准确性。
arXiv:2604.14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating t…
KPMG旗舰AI报告45处引用仅5条真实,暴露AI幻觉对商业输出的严重侵蚀。
Article URL: https://www.cityam.com/kpmg-report-on-ai-found-riddled-with-ai-hallucinations/ Comments URL: https://news.ycombinator.com/item?id=4853156…
从校准视角重新审视人机协作,揭示AI预测可靠性如何影响团队决策效率与信任。
arXiv:2606.10906v1 Announce Type: cross Abstract: We study models for human-AI teaming through the lens of statistical calibration. We assume the team…
多智能体框架实时监测并纠正医疗影像AI的模型退化和性能衰退,为临床AI可靠性保驾护航。
arXiv:2510.17004v2 Announce Type: replace-cross Abstract: Purpose: To develop and evaluate a multi-agent framework (ReclAIm) for automated monitoring,…
揭示大模型在信息验证中的致命盲点:为何AI会轻易信任却放弃核查来源?
arXiv:2606.05403v1 Announce Type: cross Abstract: Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources t…
首次系统分类MCP服务器运行时故障,揭示LLM工具化过程中的可靠性关键挑战。
arXiv:2606.05339v1 Announce Type: cross Abstract: MCP (Model Context Protocol) enables LLMs (Large Language Models) to interact with external tools an…