荣耀暂停手机地震预警服务:合作第三方出现“异常偏差”,Q4 起推送新版
IT之家 8 月 28 日消息,27 日傍晚,荣耀俱乐部团队发布《关于荣耀手机【地震预警服务】系统升级及服务维护的公告》,宣布将暂停手机地震预警服务,并于今年第三季度完成新版地震预警服务的功能适配及测试,第四季度陆续向存量机型推送系统更新。 IT之家附公告原文如下: 尊敬的荣耀用户: 近日,由于合作…
IT之家 8 月 28 日消息,27 日傍晚,荣耀俱乐部团队发布《关于荣耀手机【地震预警服务】系统升级及服务维护的公告》,宣布将暂停手机地震预警服务,并于今年第三季度完成新版地震预警服务的功能适配及测试,第四季度陆续向存量机型推送系统更新。 IT之家附公告原文如下: 尊敬的荣耀用户: 近日,由于合作…
IT之家 8 月 26 日消息,成都高新减灾研究所今晚再发声明“关于 8 月 24 日宜宾长宁 4.7 级地震的预警首报震级偏差大的声明四”。 摘要: 1、成都高新减灾研究所(以下简称减灾所)再次对 24 日宜宾长宁 4.7 级地震的首报预警震级偏差大表示道歉,但该预警不是误报; 2、减灾所未冒用“…
揭秘多模态大模型在持续学习中的公平性后门攻击,锚定偏差如何悄然埋下隐患。
arXiv:2608.21577v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in high-stakes domains where fair…
“IT早报”时间,大家好,现在是 2026 年 8 月 25 日星期二,今天的重要科技资讯有: 1. 2026 胡润中国品牌榜发布:苹果蝉联第一,抖音与微信并列,华为重返前十,AIGC 品牌首次入榜 胡润研究院发布《2026 胡润中国品牌榜》,苹果蝉联第一,AIGC 行业首次独立入榜,华为重返总榜前…
别让AI哄着你做决定!拆穿聊天机器人迎合式回答的真相,教你保持清醒判断。
Recently, I was discussing usage of AI Chatbot (chatgpt/claude) with one of my friends. She had become so dependent on AI Chatbot that almost every qu…
LLM裁判也有“偏心眼”:自我标签与他者标签会引发双向评估偏差,AI评测可信度再遭挑战。
arXiv:2608.18091v1 Announce Type: cross Abstract: As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to…
提问方式竟能左右大模型性别偏见,揭秘语言习惯背后的隐性歧视,AI公平性研究者必读
arXiv:2608.13328v1 Announce Type: cross Abstract: Professional communication is increasingly mediated by LLMs - but do these models serve all users eq…
揭秘查询时序让大模型与人类产生相反位置偏好的新发现,获奖论文值得一读。
arXiv:2608.12387v1 Announce Type: cross Abstract: Positional biases such as recency and primacy effects have been documented in large language models …
大模型竟对五条腿的狗视而不见?这项研究用溯因偏好学习破解提示不敏感难题。
arXiv:2510.09887v3 Announce Type: replace Abstract: Vision and language models frequently ignore semantically critical input edits, defaulting to pret…
IT之家 8 月 8 日消息,猫头鹰(Noctua)昨日(8 月 7 日)发布博文,指出机箱厂商标注的 CPU 散热器限高有时候并不可靠, 为此其团队手动测量超过 100 款 PC 机箱,发现 56 款与厂商规格存在相关偏差。 为解决这一问题,猫头鹰为每款已测机箱制作独立表格,标明哪些 Noctua…
揭示LLM评审在创意评估中重风格轻实质的偏见,挑战AI评价可靠性。
arXiv:2608.01666v1 Announce Type: cross Abstract: However, whether these judges truly evaluate the scientific substance of ideas or are influenced by …
揭秘人类在AI可解释性标注中的“平均偏差”,提醒研究者别盲信局部忠实性判断,值得做XAI的人细读。
arXiv:2608.00205v1 Announce Type: new Abstract: Evaluation of faithfulness of text summarization treats a model generated summary as faithful only if …
用MUD搭建AI评估场,揭露LLM裁判在聚合kappa指标下隐藏的失真,挑战传统评估框架的盲区。
Article URL: https://www.lesswrong.com/posts/GPbWyHgx9hCLMdAjc/mud-as-ai-evaluation-and-llm-judge-distortion-in-ways Comments URL: https://news.ycombi…
用反事实分析法揭示医学影像AI中的隐藏偏差,推动更可靠的AI对齐。
arXiv:2504.19621v2 Announce Type: replace Abstract: Machine learning (ML) systems for medical imaging have demonstrated remarkable diagnostic capabili…
当大模型成为历史学家:利用生成式幻觉构建历史意义,颠覆对AI时间偏差的认知。
arXiv:2607.24750v1 Announce Type: new Abstract: Large Language Models (LLMs) are temporally overexposed: trained on vast contemporary corpora, they en…
参与式设计AI偏好智能体可能引发过度信任,需警惕用户对LLM代理的盲目依赖。
arXiv:2607.21757v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate human preferences in research and practical …
医疗AI安全评估需警惕评判偏差:信息缺失下,同源评判者与人类评分员会改变结果分布,论文提出开放对话安全新标准。
arXiv:2607.18828v1 Announce Type: new Abstract: Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks. We exten…
揭示LLM多样性丧失的深层数据原因,提出Verbalized Sampling方法缓解模式坍塌
arXiv:2510.01171v4 Announce Type: replace Abstract: Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collaps…
当大模型回答问题时,其内置价值观正悄然左右输出,揭示被忽视的“价值泄漏”现象
arXiv:2607.14345v1 Announce Type: new Abstract: People use language models for practical questions whose answers are difficult to verify. We show that…
研究发现LLM评估器在不同语言中存在系统性偏见,揭示多语言场景下AI评估的公平性挑战
arXiv:2607.14480v1 Announce Type: new Abstract: LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwis…