Oura 遭集体诉讼,被指夸大智能戒指睡眠追踪准确性
IT之家 8 月 23 日消息,一项拟议中的集体诉讼正在指控智能戒指制造商 Oura 欺骗消费者,夸大其睡眠追踪功能的准确性。 这起诉讼由旧金山的 Clarkson Law Firm 于当地时间周四提起。诉状称,Oura 戒指无法测量评估睡眠质量或判断睡眠阶段所需的任何生理信号,实际上依赖 AI 生…
IT之家 8 月 23 日消息,一项拟议中的集体诉讼正在指控智能戒指制造商 Oura 欺骗消费者,夸大其睡眠追踪功能的准确性。 这起诉讼由旧金山的 Clarkson Law Firm 于当地时间周四提起。诉状称,Oura 戒指无法测量评估睡眠质量或判断睡眠阶段所需的任何生理信号,实际上依赖 AI 生…
GPT-5.6 Sol带来更准更稳的AI回答,免费用户也能用,日常对话还有Luna无限畅聊。
ChatGPT introduces improved GPT-5.6 Sol with better accuracy and consistency, plus expanded access for free users and unlimited everyday chats with GP…
实时市场数据实测,ChatGPT与Perplexity在延迟、准确度和幻觉率上分出高下。
ChatGPT vs. Perplexity: Which AI Handles Live Market Data Better? Live financial data is the ultimate stress test for LLMs. Web search latency, strict…
IT之家 7 月 26 日消息,据英国《金融时报》报道,美国阿德尔菲大学学生奥赖恩 · 纽比(Orion Newby)因被校方误指控使用 AI 生成论文、违反学术诚信原则,不得不经历一场漫长的维权过程,才最终证明自己的清白。 今年 1 月,纽约一家法院判决纽比胜诉。判决书显示,学校采用的 Turni…
如何让大模型提取的数据可信到可行动?独立验证是关键一步,解决信任难题。
Article URL: https://www.zakihsn.com/writing/ai/can-i-trust-the-numbers Comments URL: https://news.ycombinator.com/item?id=49005971 Points: 1 # Commen…
利用大语言模型低成本生成调查数据时,准确性因问题波动,本文研究如何最优分配固定的人类样本预算以提升估计效果。
arXiv:2604.17267v2 Announce Type: replace Abstract: Large Language Models can generate synthetic survey responses at low cost, but their accuracy vari…
如何让大模型更可靠地调用Web API?这篇论文提出RAG+约束解码双管齐下,显著削减调用错误,实战价值高。
arXiv:2607.05936v1 Announce Type: cross Abstract: Integration of web APIs is a cornerstone of modern software systems, yet writing correct web API inv…
用Python脚本同时向5个AI检测器发送同一篇文章,82%与31%的巨大差异揭露检测工具的可信度危机。
Last semester I pasted the same essay into two detectors back to back and got 82% on one, 31% on the other. No edits. Same paste. Within the same hour…
GPT-5挑战Scrum认证考试,实证准确率研究新发现
arXiv:2607.00049v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in Agile Software Development for documentation, …
揭秘大模型校准排名反转现象,提出准确性控制评估框架,让LLM比较更公平。
arXiv:2606.30814v1 Announce Type: new Abstract: Calibration evaluates whether a model confidence aligns with its empirical accuracy. Existing studies …
面对LLM时代的软件度量困境,该帖引发关于准确性、成本、上下文等多维度指标如何平衡的深度思考
A bit of a rant. Sorry! With the probablistic pluggable 'brain' existing in parts of the solution how are you measuring anything is better or worse? I…
AI深度学习模型仅凭非对比MRI即可预测脑肿瘤强化,多队列研究验证高准确性,有望减少对比剂使用风险
arXiv:2508.16650v3 Announce Type: replace-cross Abstract: Brain tumour MRI typically requires both pre- and post-contrast imaging, but gadolinium is n…
首次聚焦多轮LLM对话中非功能性需求评估,揭示准确性与用户满意度的新挑战
arXiv:2606.24834v1 Announce Type: new Abstract: LLM-based dialogue assistants have become mainstream tools for software developers, yet current evalua…
探讨蒙特卡洛树搜索在开源AI编码工具中提升代码生成准确性的可能性,问题导向的讨论帖
Do you consider that Monte Carlo Tree Search implementation in an opensource environment (like codex, claude code) for AI can improve the code generat…
大模型能否胜任医生角色的关键考验:最新研究实探LLM在医疗诊断与临床推理评分中的准确性。
arXiv:2604.14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating t…
AI记忆工具可能适得其反,新研究揭示它们如何让模型回答更差而非更好,引发对用户偏好与准确性的权衡思考。
New research suggests that AI memory systems can degrade model performance and encourage sycophantic tendencies.
深度伪造检测新突破:兼顾校准、公平与准确性,为AI安全提供可靠新方法。
arXiv:2606.09881v1 Announce Type: new Abstract: Deepfake detectors show large performance gaps across demographic groups. Existing fairness approaches…
当大模型没把握时,与其放弃回答,不如学会“模糊处理”——一种选择性抽象策略提升长文本生成可靠性
arXiv:2602.11908v3 Announce Type: replace Abstract: LLMs are widely used, yet they remain prone to factual errors that erode user trust and limit adop…
AI智能体的可信度依赖数据根基:再聪明的模型,错数据也会导致决策风险。
Your finance team built an agent that helps close the books. It connects to the ERP, reads journal entries, and drafts reconciliations. In the demo, e…
临床LLM对同一患者表述语义变化可能导致诊断不一致,稳定性亟待关注。
arXiv:2605.30646v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in clinical applications. However, their behavior…