谷歌推出全球首个双盲 AI 评估技术,基于加密环境保障基准测试公正性
IT之家 8 月 27 日消息,谷歌宣布推出 全球首个针对前沿专有 AI 模型的双盲评估 ,将外部评估限制在加密“黑箱”环境中,避免模型提前获取测试信息来优化性能。 就像学生考前不能提前看到试题才能真实反映水平一样,当前 AI 模型评估也面临同样问题:如果模型提前接触测试题目(IT之家注:也就是所谓…
IT之家 8 月 27 日消息,谷歌宣布推出 全球首个针对前沿专有 AI 模型的双盲评估 ,将外部评估限制在加密“黑箱”环境中,避免模型提前获取测试信息来优化性能。 就像学生考前不能提前看到试题才能真实反映水平一样,当前 AI 模型评估也面临同样问题:如果模型提前接触测试题目(IT之家注:也就是所谓…
破解AI看病难题:全新基准MedReaMM如何考核多模态大模型的专业诊断整合能力
arXiv:2608.22323v1 Announce Type: new Abstract: The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing int…
大模型写代码常“编造”不存在依赖包?这篇论文系统评估了推理时防御手段,帮你避开代码幻觉坑。
arXiv:2608.22652v1 Announce Type: cross Abstract: LLMs are increasingly used for code generation, yet they frequently hallucinate non-existent softwar…
用粤语语法资源当试金石,实测大模型做知识驱动语法工程的成色,实验设计严谨
arXiv:2608.23448v1 Announce Type: new Abstract: This paper presents new Cantonese ParGram resources and evaluates LLMs for knowledge-driven grammar en…
打破问卷、选择与生成文本混为一谈的评测假设,为LLM价值测量提供分层契约验证。
arXiv:2608.23411v1 Announce Type: new Abstract: LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from genera…
揭秘LLM路由差距的根源:任务类型比模型选择更关键,实测14模型两次运行结果并不完全一致。
arXiv:2608.23023v1 Announce Type: new Abstract: An LLM router picks which model should answer each query. The appeal is that models fail on different …
当单个大模型评代码不够可信,让多个LLM组成「陪审团」一致表决会怎样?实验数据揭示真相
arXiv:2602.18492v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are now good enough at coding that developers can describe inte…
把真实事故变成考试题,比凑数量更有用,手把手教你将这份评估模板移植到团队流程中。
So far this series has been about giving my order-reading LLM an exam . Some of you have been reading it thinking: "Mine isn't orders, it's meeting-mi…
LLM辅助写作效果究竟如何?这篇论文呼吁用实证方法检验,而非凭感觉判断。
arXiv:2608.22124v1 Announce Type: new Abstract: LLM-assisted writing is often treated as a detection problem, as it raises questions about clarity, in…
用数字孪生搭建“诊所”,让具身AI在真实任务中接受标准化体检,评估更靠谱。
arXiv:2608.21416v1 Announce Type: cross Abstract: Embodied artificial intelligence (AI) must be tested in the clinical environments where it will oper…
首个4K级AI生成图像缺陷检测数据集,支持检测、定位与解释三合一,是图像生成质量评估的利器。
arXiv:2608.20713v1 Announce Type: new Abstract: Generative AI can now produce highly realistic images, yet current models still exhibit subtle but cri…
科研新基准:从发明人式披露生成专利文本,直击大模型专业写作能力短板。
arXiv:2608.21249v1 Announce Type: new Abstract: While recent large language models (LLMs) have achieved promising results on individual patent draftin…
药物发现遇上智能体,看LLM评估系统如何借人类对齐把关可靠性
arXiv:2608.21057v1 Announce Type: new Abstract: Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug di…
揭秘专用裁判与共享评判的取舍之道,为LLM评估定制最优策略,实验详实、图表清晰。
arXiv:2607.27984v2 Announce Type: replace Abstract: Agentic systems generate outputs faster than human review. We contrast two LLM evaluator specializ…
用LLM当裁判评估大模型在5G故障诊断中的表现,为垂直领域落地提供新思路。
arXiv:2608.21021v1 Announce Type: cross Abstract: Real-world fault analysis in 5G and emerging 6G networks demands domain expertise to analyze free-te…
大模型当裁判时,信任评分与事实判断真能独立吗?这项研究用受控QA测试揭开了两者的纠缠关系
arXiv:2608.21097v1 Announce Type: new Abstract: LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, …
人机分歧竟是金矿?用LLM与人类判断差异改进质量评估,论文方法很新颖。
arXiv:2608.20385v1 Announce Type: new Abstract: Systematic reviews rely on quality appraisal of included studies, a process that is time-consuming and…
变异测试一招揪出LLM评测漏洞,Muteval让评估数据更可靠,开源免费直接上手。
Article URL: https://github.com/AshwinUgale/muteval Comments URL: https://news.ycombinator.com/item?id=49388918 Points: 1 # Comments: 0
前沿AI模型的政治立场、伦理取向与人格特质,一次基准测试看透底牌
Article URL: https://www.blackbench.ai/ Comments URL: https://news.ycombinator.com/item?id=49382540 Points: 3 # Comments: 2
安全专家实测122次AI攻防挑战,发现10次违规行为,揭示大模型在安全任务中的失控风险。
The AI Security Institute has a new report of AI systems engaging in “unsanctioned behavior”—what I have been calling “ genie …