Your brain on AI
用AI帮你快速验证新闻真伪,但别忘了保留独立判断。Perplexity AI提供实时信息检索与来源引用,边查边学,适合对抗假新闻。
Many people find AI-based chatbots helpful in keeping up with news, but a study by Pattie Maes and her colleagues at the MIT Media Lab points to a big…
用AI帮你快速验证新闻真伪,但别忘了保留独立判断。Perplexity AI提供实时信息检索与来源引用,边查边学,适合对抗假新闻。
Many people find AI-based chatbots helpful in keeping up with news, but a study by Pattie Maes and her colleagues at the MIT Media Lab points to a big…
大模型当裁判时,信任评分与事实判断真能独立吗?这项研究用受控QA测试揭开了两者的纠缠关系
arXiv:2608.21097v1 Announce Type: new Abstract: LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, …
大模型微调后“知道却说不出口”的难题,用召回锚定蒸馏可精准破解。
arXiv:2608.20794v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) can degrade factual behavior outside the target domain. This degradation …
代码助手RAG分不清新旧事实?这篇研究用真实软件历史消除陈旧事实错误,直击AI记忆的时间盲区。
arXiv:2608.20685v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) has no model of time: when a fact changes across a coding sessi…
同一问题换个问法,大模型对事实与信念的判断就不同,这份研究揭示了提示词如何影响模型的认知边界。
arXiv:2608.17809v1 Announce Type: new Abstract: Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "…
让ChatGPT告别瞎编,查询真实世界地理事实的确定性插件,LLM与空间数据结合的玩法值得一试。
Try it in ChatGPT: @emem What has changed around this location? Or ask it something about our world. It is the perfect blend of non-deterministic LLM …
以色列疑似伪造智库数据集,试图暗中操纵ChatGPT等AI的回答,揭露AI信息战的全新战线。
Article URL: https://responsiblestatecraft.org/israel-influence-chatgpt/ Comments URL: https://news.ycombinator.com/item?id=49337392 Points: 41 # Comm…
AI研究速览:聚焦评估污染与检索扩展,一窥模型评测新盲区。
Reliable provenance and graded trust Explicit reliability modeling cuts hallucination. Σ‑Mem stores symmetric competence states for peer agents, while…
把AI记忆锚定在可重算的事实上,让主观判断有据可依,轻量开源新思路。
Article URL: https://github.com/Anchorstate-Lab/GMR Comments URL: https://news.ycombinator.com/item?id=49292758 Points: 1 # Comments: 1
研究大模型在事实性主张上的分歧,帮你看清不同LLM的可靠性差异与潜在盲点。
Article URL: https://zenodo.org/records/21829261 Comments URL: https://news.ycombinator.com/item?id=49272730 Points: 1 # Comments: 1
用反事实基准和训练方法破解大模型事实一致性与异质知识推理的排序难题
arXiv:2608.07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided…
黑盒大模型也能验明正身?反事实指纹技术为LLM所有权验证提供新思路,防窃取。
arXiv:2608.08195v1 Announce Type: cross Abstract: Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tu…
用因果推断给数字孪生做“证伪测试”,让仿真模型不再自说自话。
arXiv:2301.07210v5 Announce Type: replace-cross Abstract: Digital twins are simulation-based models designed to predict how a real-world process will …
用反事实模拟训练提升思维链忠实度,解读大模型推理黑箱的新思路。
arXiv:2602.20710v2 Announce Type: replace Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM p…
探究推理链长度如何影响大模型对答案事实性的判断,揭示关键机制。
arXiv:2604.06756v2 Announce Type: replace Abstract: Large language models (LLMs) has been widely adopted as a scalable surrogate for human evaluation,…
一句话看懂LLM幻觉检测新框架,轻量黑盒免参考,不止摘要还跨任务实测。
arXiv:2608.05823v1 Announce Type: new Abstract: The reliability of Large Language Models (LLMs) is often compromised by factual inconsistencies, inclu…
用GPT-3.5预测反事实结果,在线借贷场景下的决策新思路。
arXiv:2608.05367v1 Announce Type: new Abstract: Counterfactual analysis aims to predict potential outcomes under hypothetical scenarios, offering valu…
大模型如何在激活空间内编码“上下文真值”?这项研究揭示了线性方向的全新机制,带你深入AI判断真相的内部逻辑。
arXiv:2608.03035v1 Announce Type: new Abstract: Prior work has shown that LLMs encode the truth of factual propositions along linear directions in act…
用模拟器增强检索,让大模型在长科学问答中更精准地锚定事实,RAG落地新思路。
arXiv:2509.25459v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) show promise in generating long-form scientific explanations th…
用图思维链改进因果发现,并揭示事后路径公平审计的脆弱性,适合关注AI公平与因果推理的研究者。
arXiv:2608.02877v1 Announce Type: new Abstract: Causal discovery recovers directed structure from observational data and is increasingly used in clini…