Improved Confidence Estimates for Black-Box Large Language Models
黑盒大模型置信度不准?这项研究提出新方法,让答案可信度更可靠。
arXiv:2608.19323v1 Announce Type: new Abstract: Uncertainty quantification (UQ) is essential for the safe deployment of large language models (LLMs). …
黑盒大模型置信度不准?这项研究提出新方法,让答案可信度更可靠。
arXiv:2608.19323v1 Announce Type: new Abstract: Uncertainty quantification (UQ) is essential for the safe deployment of large language models (LLMs). …
让大模型学会“不知道就承认”,用激励机制打造可靠的选择性回答,直击幻觉痛点。
arXiv:2604.03904v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often produce confident but incorrect answers, in part because …
探究大模型在代码补全中的自信与准确之谜,揭示“愚者确定,智者存疑”的真相。
arXiv:2508.16131v3 Announce Type: replace-cross Abstract: Code completion entails the task of providing missing tokens given a surrounding context. It…
不看全文也能“有把握放弃回答”?这项研究为选择性问答系统提供了风险校准的理论保障。
arXiv:2608.12008v1 Announce Type: new Abstract: Large language models (LLMs) may generate fluent but incorrect answers, making uncertainty quantificat…
医疗视觉问答也能“知道何时不懂”,置信度感知推理让AI诊断更可靠
arXiv:2608.10964v1 Announce Type: new Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produc…
LLM不确定性新思路:无需额外训练,仅靠注意力路径的脆弱程度就能暴露模型“心虚”时刻,极具实用潜力。
arXiv:2608.11138v1 Announce Type: new Abstract: We propose that a model's uncertainty about a token is reflected not only in the breadth of its output…
开源大模型以细粒度辩论实现可信数据扩充,为心理健康与网络内容安全标注开辟新路径。
arXiv:2512.06227v3 Announce Type: replace-cross Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) appli…
揭秘LLM推荐系统幻觉:模型是否自知?联合审计幻觉率与置信度校准,为目录忠实度提供新视角。
arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-$K$ item suggestion regularly emit titles outside the target catalog. Prior…
用测试时训练校准大模型推理置信度,让 conformal 方法在分布变化下依然可靠,COLM 2026 新思路值得关注。
arXiv:2604.01170v2 Announce Type: replace-cross Abstract: While test-time scaling has enabled large language models to solve highly difficult tasks, s…
大模型越自信越容易说谎?这篇论文揭示置信度如何放大LLM的欺骗风险,对AI安全研究有重要启示。
arXiv:2607.20444v1 Announce Type: cross Abstract: Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service …
提出概率置信度选择与排序方法,优化大模型推理链,提升复杂推理的准确性与可解释性。
arXiv:2508.21787v3 Announce Type: replace-cross Abstract: Best-of-n sampling improves the accuracy of large language models (LLMs) and large reasoning…
物理运动感知与置信度引导LLM结合的生成推理新方法,被ECCV 2026接收
arXiv:2505.16456v3 Announce Type: replace Abstract: Recent advances in 3D content generation have amplified demand for dynamic models that are both vi…
一句话锁定大模型生成置信度的新方法,粒度精细到片段级,告别粗粒度评估。
arXiv:2607.05721v1 Announce Type: new Abstract: Uncertainty estimation is essential not only for the trustworthy deployment of large language models (…
系统解剖LLM中不确定性的来源与量化方法,为可信AI提供理论基础。
arXiv:2603.24967v2 Announce Type: replace Abstract: Understanding why a large language model (LLM) is uncertain about the response is important for th…
开源项目Aletheia探索无测试集下的循环工程,通过明确不确定性视图和矛盾证据降低置信度,为Claude Code和Codex注入推理智能。
Aletheia is an open-source attempt to explore what loop engineering looks like when reality has no test suite. It maintains an explicit view of what m…
最新论文提出Fork-Think方法,通过分叉思考增强置信度,提升推理可靠性。
arXiv:2606.31484v1 Announce Type: new Abstract: Parallel thinking has enjoyed great success for boosting LLM performance on reasoning tasks without th…
LLM说“我很有把握”其实更代表“我坚持选它”,而非正确答案——揭秘AI置信度的真实含义。
arXiv:2606.29490v1 Announce Type: new Abstract: Confidence is an estimate of the probability that a chosen answer is correct. Verbal confidence report…
研究修剪注意力层如何影响LLM的解释忠实性与置信度校准,揭示模型优化新视角。
arXiv:2606.24970v1 Announce Type: new Abstract: Pruning Large Language Models (LLMs) reduces memory and inference costs by removing parts of the netwo…
基于置信度和提示注入检测,为OpenAI智能体行动加一道安全门,60秒本地部署。
Article URL: https://github.com/Lelu-ai/lelu Comments URL: https://news.ycombinator.com/item?id=48664025 Points: 4 # Comments: 0
大语言模型看似自信输出时,内部却在处理认知失调,这篇论文揭示了这种矛盾心理机制。
arXiv:2606.22633v1 Announce Type: new Abstract: Large language models (LLMs) frequently encounter inputs that disagree with their prior outputs, throu…