AI Weekly: Four Frontier Models in Four Days
四天连发四款前沿大模型,Claude与Gemini等新王对决,基准成绩深度解析,AI圈一周动态全掌握。
Week of August 11 to 18, 2026 Four labs shipped frontier models within four days of each other this week, and every one of them was tuned for the same…
四天连发四款前沿大模型,Claude与Gemini等新王对决,基准成绩深度解析,AI圈一周动态全掌握。
Week of August 11 to 18, 2026 Four labs shipped frontier models within four days of each other this week, and every one of them was tuned for the same…
Artificial Analysis推出LLM搜索API对比评测,900个任务F1打分,帮你挑选最佳检索模型。
Article URL: https://artificialanalysis.ai/agents/search-api Comments URL: https://news.ycombinator.com/item?id=49359449 Points: 1 # Comments: 0
分布式系统bug修复遇上AI智能体,DDBench双条件评测揭示其真实能力边界。
arXiv:2608.14863v1 Announce Type: cross Abstract: LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now …
首个系统评估多模态大模型生成图表注释能力的基准,揭示模型在数据推理与视觉布局上的短板。
arXiv:2608.03464v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have made significant progress in chart understanding, gene…
首个统一基准横评多任务3D脑肿瘤分割模型,MRI医学影像AI的实用参考。
arXiv:2607.28858v1 Announce Type: cross Abstract: Automatic brain tumor segmentation from magnetic resonance imaging (MRI) has become a fundamental ta…
AI编程基准评测常被作弊污染,Proctor用签名隔离包让作弊无处遁形,值得一看。
Article URL: https://github.com/dylanp12/proctor Comments URL: https://news.ycombinator.com/item?id=48650355 Points: 2 # Comments: 0
让AI当裁判,检验它能否分辨人类与机器对话,反图灵测试新基准。
arXiv:2606.21844v1 Announce Type: new Abstract: As AI systems integrate into online spaces, differentiating them from humans in conversations is incre…
多款大模型同台竞技预测2026世界杯,排名揭晓谁最懂足球
Article URL: https://llmsoccerarena.up.railway.app/ Comments URL: https://news.ycombinator.com/item?id=48538417 Points: 2 # Comments: 1
LLM编码代理如何优化大型代码库?这篇论文提出FormulaCode基准,评估真实场景下的整体优化能力,超越传统合成任务与二值信号。
arXiv:2603.16011v2 Announce Type: replace-cross Abstract: Large language model (LLM) coding agents increasingly operate at the repository level, motiv…