LLM Reasoning Traces Are Not Audit Records
大模型的思维链并非真实决策的审计日志,揭开推理可解释性的关键盲区。
Article URL: https://rye.ai/blog/cot-faithfulness-reasoning-traces-not-audit-logs/ Comments URL: https://news.ycombinator.com/item?id=49371034 Points:…
大模型的思维链并非真实决策的审计日志,揭开推理可解释性的关键盲区。
Article URL: https://rye.ai/blog/cot-faithfulness-reasoning-traces-not-audit-logs/ Comments URL: https://news.ycombinator.com/item?id=49371034 Points:…
AI代理失控事件后,OpenAI加码安全监控,揭秘思维链审查新机制。
The ChatGPT maker says its upcoming Astra model may have reached “critical” cyber capabilities, prompting it to halt a significant number of training …
把医学鉴别诊断逻辑引入多智能体协作,为复杂推理提供新思路。
arXiv:2608.11420v1 Announce Type: new Abstract: Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications …
揭示思维链并非万能,实证剖析串行深度瓶颈对LLM推理的边界影响。
arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We in…
新技巧让AI的“内心独白”无所遁形,还疑似抓到模型间蒸馏痕迹
Researchers devised a way to extract “reasoning traces” from Claude, GPT, and Gemini. What they found, they say, indicates that some Chinese AI may be…
用稀疏自编码器拆解大模型思考与直接回答的神经差异,揭开推理黑箱的机制奥秘。
arXiv:2608.08168v1 Announce Type: new Abstract: While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabil…
用反事实模拟训练提升思维链忠实度,解读大模型推理黑箱的新思路。
arXiv:2602.20710v2 Announce Type: replace Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM p…
推理token竟分“结构”与“内容”两类,熵引导超令牌可压缩思维链,直击大模型推理成本痛点。
arXiv:2604.26355v4 Announce Type: replace Abstract: Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level …
用平均场理论拆解思维链推理的动态机制,为理解大模型推理过程提供全新数学视角
arXiv:2608.05152v1 Announce Type: cross Abstract: Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent year…
百万参数小模型也能用思维链推理?这项研究让推理机制分析不再只属于大模型。
arXiv:2608.04980v1 Announce Type: cross Abstract: We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we cal…
从思维链的动态变化中识别大模型推理失误,为AI安全与可解释性提供全新监测路径,值得关注。
arXiv:2608.03291v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing …
1.58比特量化新框架ScaleQ-1.58,专治推理大模型量化掉点,用三元PTQ保住思维链能力,硬核干货。
arXiv:2608.01078v1 Announce Type: cross Abstract: We propose ScaleQ-1.58, a scalable ternary post-training quantization (PTQ) framework for reasoning …
用“推理能量”量化思维链每步开销,为LLM思考效率提供新视角
arXiv:2607.28674v1 Announce Type: new Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning…
用思维链蒸馏提升表格重排序效果,TabRank创新方法亮相。
arXiv:2607.25182v1 Announce Type: cross Abstract: The ability to retrieve relevant tables for answering questions is a key task for structured informa…
通过验证器规模化合成长期思维链,大幅提升LLM在数学和编程等场景的推理能力。
arXiv:2509.03059v2 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can b…
这篇论文揭示了大模型推理过程中有大量隐蔽计算并未呈现在思维链中,挑战了我们对Chain-of-Thought透明性的固有认知。
arXiv:2607.22925v1 Announce Type: cross Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its outpu…
破解AI推理中的病态思维链,这篇论文为提升大模型可靠性提供诊断新方法。
arXiv:2602.13904v2 Announce Type: replace Abstract: Chain-of-thought (CoT) reasoning is fundamental to modern LLM architectures and represents a criti…
AI安全里程碑:GPT-RED发现人类从未见过的「假思维链」攻击,AI攻防进入自我博弈时代。
针对长思维链大语言模型的KV缓存量化新突破,渐进混合精度方案,已被ICLR 2026接收,代码已开源。
arXiv:2505.18610v2 Announce Type: replace Abstract: Recently, significant progress has been made in developing reasoning-capable Large Language Models…
提出概率置信度选择与排序方法,优化大模型推理链,提升复杂推理的准确性与可解释性。
arXiv:2508.21787v3 Announce Type: replace-cross Abstract: Best-of-n sampling improves the accuracy of large language models (LLMs) and large reasoning…