The Fragility of Strategic Thinking in Large Language Models
不只是问答,这项研究揭示了LLM在战略博弈中的推理短板,值得关注。
arXiv:2510.10813v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly applied to domains that require reasoning about othe…
不只是问答,这项研究揭示了LLM在战略博弈中的推理短板,值得关注。
arXiv:2510.10813v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly applied to domains that require reasoning about othe…
探究大模型能否识别自身易错场景,为AI安全与可靠部署提供新视角。
arXiv:2607.23496v2 Announce Type: replace Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the sam…
用图思维链改进因果发现,并揭示事后路径公平审计的脆弱性,适合关注AI公平与因果推理的研究者。
arXiv:2608.02877v1 Announce Type: new Abstract: Causal discovery recovers directed structure from observational data and is increasingly used in clini…
AI对齐不完美时,价值系统有多脆弱?这篇论文用严谨分析揭示潜在风险,值得关注。
arXiv:2607.28881v1 Announce Type: new Abstract: As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that …
视觉Transformer部署提速新思路,用脆弱性引导混合精度量化,兼顾压缩率与精度的双赢方案。
arXiv:2607.28589v1 Announce Type: cross Abstract: Post-training quantization (PTQ) has emerged as an effective solution for deploying Vision Transform…
揭秘大模型如何被输入顺序“带偏”——这项研究深入分析了LLM偏好的脆弱性,对理解模型鲁棒性至关重要。
arXiv:2506.14092v4 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in decision-support systems for high-stakes…
从成功流程中追溯智能体失败根源,揭示AI自主决策的脆弱性
arXiv:2607.12747v1 Announce Type: cross Abstract: Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajec…
揭示语言模型自我生成QA训练的隐藏脆弱性,一篇值得关注的AI研究论文。
arXiv:2606.32002v1 Announce Type: new Abstract: Language models are increasingly taught from synthetic question--answer (QA) supervision: a model gene…
揭示AI模型在静默数据损坏下的脆弱性,为可靠性部署提供关键参考。
arXiv:2405.01741v4 Announce Type: replace-cross Abstract: Reliability of AI systems is a fundamental concern for the successful deployment and widespr…
当AI悄悄主导对话走向,你可能正陷入共构盲区却不自知——这篇CHI论文直击人机协作最隐秘的认知风险。
arXiv:2606.20762v1 Announce Type: cross Abstract: This paper introduces two constructs to describe, as far as we know, a previously unnamed risk in hu…
揭秘LLM安全评估的致命缺陷:看似全面防护实则暗藏针对特定群体的盲区,戳破“选择性安全陷阱”系统性风险
arXiv:2601.04389v3 Announce Type: replace-cross Abstract: Current safety evaluations of large language models (LLMs) create a dangerous illusion of un…
用几何方法诊断大模型对齐脆弱性,揭示少量微调即可抹除安全拒答的隐患
arXiv:2606.22676v1 Announce Type: new Abstract: Alignment tuning is meant to make harmful-request refusal robust, yet this safety behavior can be eras…
揭秘无训练AI图像检测器的脆弱性,系统评估分数方向、预处理与压缩的三重影响。
arXiv:2606.20488v1 Announce Type: new Abstract: Training-free detectors of AI-generated images promise generator-agnostic deployment without classifie…
论文揭示防御训练会让LLM智能体付出“自主性税”,性能与安全如何平衡?
arXiv:2603.19423v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly rely on external tools (file operations, API …
迁移学习解决结构脆弱性建模中数据稀缺难题,方法首创且案例充分。
arXiv:2606.18567v1 Announce Type: cross Abstract: This paper presents a methodology-centered transfer learning framework for fragility adaptation unde…
借助伪提示语言(Pseudo Prompting Language),系统化解决自然语言提示词中角色、目标与约束松散导致的交互脆弱性,提升生成式AI与智能体对齐效率。
arXiv:2606.17164v1 Announce Type: cross Abstract: Prompting has become the primary interface between humans and generative AI, yet many natural langua…
这篇arXiv论文提出“认知债务”概念,剖析AI作为智力杠杆如何引发系统性脆弱性,值得技术决策者深思。
arXiv:2606.15078v1 Announce Type: new Abstract: We develop a formal theory of cognitive debt: the stock of unverified reasoning obligations that accum…
音频模型解释并非可信:研究发现可在不改变预测结果的情况下操纵归因,揭示AI可解释性的潜在脆弱性。
arXiv:2606.14466v1 Announce Type: cross Abstract: This paper investigates the fragility of post-hoc explanation methods in audio deepfake detection. W…
揭示医学影像AI在肺部结节检测中受采集状态影响的内在不稳定性,为AI治理提供定量依据
arXiv:2606.12824v1 Announce Type: cross Abstract: AI governance for medical imaging is formalizing: the 2026 ACR-SIIM Practice Parameter recommends lo…
当标准探测准确率饱和时,引入“脆弱性”度量作为互补指标,为LLM预训练分析提供新视角。
arXiv:2606.11375v1 Announce Type: cross Abstract: Standard linear probing declares a property "encoded" when a classifier on hidden states achieves hi…