谷歌推出全球首个双盲 AI 评估技术,基于加密环境保障基准测试公正性
IT之家 8 月 27 日消息,谷歌宣布推出 全球首个针对前沿专有 AI 模型的双盲评估 ,将外部评估限制在加密“黑箱”环境中,避免模型提前获取测试信息来优化性能。 就像学生考前不能提前看到试题才能真实反映水平一样,当前 AI 模型评估也面临同样问题:如果模型提前接触测试题目(IT之家注:也就是所谓…
IT之家 8 月 27 日消息,谷歌宣布推出 全球首个针对前沿专有 AI 模型的双盲评估 ,将外部评估限制在加密“黑箱”环境中,避免模型提前获取测试信息来优化性能。 就像学生考前不能提前看到试题才能真实反映水平一样,当前 AI 模型评估也面临同样问题:如果模型提前接触测试题目(IT之家注:也就是所谓…
OpenAI自研Jalapeño芯片基准测试曝光,专为大规模快速推理而生,直面英伟达Blackwell。
Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently availa…
破解AI看病难题:全新基准MedReaMM如何考核多模态大模型的专业诊断整合能力
arXiv:2608.22323v1 Announce Type: new Abstract: The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing int…
为Claude Code打造防错技能,经591次盲测验证,六类模型均显著提升。
Article URL: https://github.com/rainmanjam/poka-yoke Comments URL: https://news.ycombinator.com/item?id=49427559 Points: 1 # Comments: 0
多模态大模型不止看懂RGB,高光谱成像理解迎来全新基准与免训练推理框架。
arXiv:2604.08884v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on RGB image under…
系统检验机器遗忘算法在极端压力下的鲁棒性,为隐私保护研究划出新基准。
arXiv:2608.22527v1 Announce Type: new Abstract: Recently, machine unlearning, the removal of specific training data influence from a model, has gained…
科研新基准:从发明人式披露生成专利文本,直击大模型专业写作能力短板。
arXiv:2608.21249v1 Announce Type: new Abstract: While recent large language models (LLMs) have achieved promising results on individual patent draftin…
前沿AI模型的政治立场、伦理取向与人格特质,一次基准测试看透底牌
Article URL: https://www.blackbench.ai/ Comments URL: https://news.ycombinator.com/item?id=49382540 Points: 3 # Comments: 2
不看任务准确率,专测AI模仿你文风的“神似度”,PersonalBench填补了个性化写作评测空白
arXiv:2608.19746v1 Announce Type: new Abstract: Personalized text generation aims to make LLMs write in a specific individual's style, yet existing be…
首个面向递归自我改进的AI智能体算法设计基准,直击AI自我进化核心命题。
arXiv:2608.20318v1 Announce Type: cross Abstract: Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI …
大模型记忆中的认知陷阱有了专门的评测基准,帮你识别AI“想当然”的隐患。
arXiv:2608.20202v1 Announce Type: cross Abstract: Memory has become a key component of large language models, enabling them to retain information and …
高通紧急修改骁龙C处理器宣传幻灯片,悄然删除两项基准测试结果,背后有何考量?
IT之家 8 月 20 日消息,科技媒体 Tom's Hardware 昨日(8 月 19 日)发布博文,报道称高通已联系多家媒体,要求改用新版骁龙 C 处理器(Snapdragon C)幻灯片, 主要删除了“空闲应用”和“网页浏览”2 项基准测试结果。 IT之家曾于 8 月 13 日报道, 高通披…
首个面向通用AI助手的智能体视频理解基准,测试AI在真实视频场景中的决策与交互能力。
arXiv:2608.14718v1 Announce Type: new Abstract: Video understanding is a fundamental task for evaluating the capabilities of multimodal large language…
重新定义Agentic-SQL,从自主性视角构建分类体系并做基准实测,值得AI研发者细读。
arXiv:2608.15389v1 Announce Type: new Abstract: LLM-based Text-to-SQL progress is reported across heterogeneous benchmarks, backbones, and inference p…
Anthropic神秘Model 2曝光,基准测试略胜当前最强Claude Mythos 5,综合能力再升级。
IT之家 8 月 18 日消息, Anthropic 于 8 月 15 日披露《2026 年 8 月风险报告》 ,其中提及尚未发布、名为 Model 2 的 AI 模型, 在多项任务上要优于目前公开的最强模型 Claude Mythos 5。 根据报告披露的细节,相比较 Claude Mythos …
系统评估五种LLM提示交付方式,用辅助不确定性信号提升系统综述筛查的可靠性与效率
arXiv:2608.14551v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for title-abstract screening in systematic review…
首个衡量LLM驱动代码编辑安全漂移的基准,揭示AI改码时隐藏的安全退化风险。
arXiv:2608.15092v1 Announce Type: cross Abstract: In this work, we introduce WeSCE, a benchmark for quantifying security drift in code editing under w…
LLM能像人类一样构造特殊图吗?这个数据集专门来测,为图论+AI交叉研究提供新基准。
arXiv:2608.14999v1 Announce Type: cross Abstract: Constructing special graphs is an important task within graph theory and computer science. Many popu…
消费级显卡也能跑“Opus级”Agent,Qwen3.8-27B多项基准反超Claude,开源小模型性能炸场。
推理能力还能自定义
更多次搜索竟比更强搜索模型更有效?OpenRouter基准测试揭示LLM搜索策略的真相。
Article URL: https://openrouter.ai/blog/announcements/web-search-benchmark/ Comments URL: https://news.ycombinator.com/item?id=49302033 Points: 2 # Co…