I benched frontier AI LLM models on politics, ethics, and personality traits
前沿AI模型的政治立场、伦理取向与人格特质,一次基准测试看透底牌
Article URL: https://www.blackbench.ai/ Comments URL: https://news.ycombinator.com/item?id=49382540 Points: 3 # Comments: 2
前沿AI模型的政治立场、伦理取向与人格特质,一次基准测试看透底牌
Article URL: https://www.blackbench.ai/ Comments URL: https://news.ycombinator.com/item?id=49382540 Points: 3 # Comments: 2
Grok 4.6正式亮相,性能超越Kimi K3、紧咬GPT-5.6 Sol,还以不到半价的API成本抢占市场。AI模型混战再升级,值得关注。
Elon Musk's company SpaceXAI, formerly known as xAI, has released Grok 4.6 , its latest frontier AI model, with a focus on long-running agents, c…
一键追踪2400+AI基准测试,快速查看模型能力排名与评分。
We’re launching BenchmarkList: one place to track AI benchmarks, models, and capabilities. It’s surprisingly hard to get a complete picture of what AI…
OpenAI揭秘热门编码基准测试的缺陷,重新审视AI模型评估的可靠性
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating …
聚焦AI癌症检测模型的基准测试,揭露人口统计与扫描协议偏差如何影响诊断性能,为医疗AI公平性提供关键参考。
arXiv:2606.24883v1 Announce Type: new Abstract: Artificial intelligence (AI) has achieved remarkable success in medical imaging, but it is widely reco…
网络设备配置翻译迎来可复现基准,多厂商DSM转CLI的语义评测新标准,用AI打通状态模型与命令行鸿沟。
arXiv:2606.20564v1 Announce Type: cross Abstract: Translating high-level network intents into correct multivendor configurations remains a central cha…
首个评估AI网页智能体在网站安全与隐私任务上的基准,揭示现有模型的脆弱性
arXiv:2604.06367v2 Announce Type: replace-cross Abstract: Web agents automate browser tasks, ranging from simple form completion to complex workflows …
首个超分子化学AI基准测试,推动分子计算与新药研发
arXiv:2606.13477v1 Announce Type: cross Abstract: Supramolecular chemistry, which includes the study of non-covalent host-guest assemblies, has advanc…
AI基准测试忽略的致命短板:买够GPU却卡在数据路径的突发瓶颈上。
Presented by F5 Enterprise AI teams have spent years solving for compute, securing GPU allocations, negotiating cloud capacity, and benchmarking train…
大型语言模型能否像城市规划师一样推理?研究用专业判断基准测试AI,揭示AI与人类专家的能力边界。
arXiv:2606.11678v1 Announce Type: new Abstract: Problem, Research Strategy, and Findings: The rise of large language models (LLMs) raises a key questi…
最新基准测试显示,Anthropic的Claude模型在抵制俄罗斯宣传方面表现最优,Opus 4.7获得“杰出”评级。
Estonian government benchmark shows how dozens of models combat Russia's "strategic narratives."
首个专门评估LLM辅助同行评审行为的基准,揭示AI评审的可靠性边界。
arXiv:2605.29815v1 Announce Type: new Abstract: The growing number of submitted papers has motivated the exploration of Large Language Models (LLMs) a…
揭示AI排行榜得分背后的测量噪声,并提出框架量化潜在能力差异,颠覆你对基准测试的认知。
arXiv:2605.25272v1 Announce Type: new Abstract: While aggregate leaderboard scores drive AI development, they contain substantial measurement noise wh…
AI基准测试暗藏理论假设,窄化进步定义,警惕评估陷阱重塑能力概念
arXiv:2605.14167v1 Announce Type: new Abstract: Every AI benchmark operationalizes theoretical assumptions about the capability it claims to assess. W…
看AI如何像人类一样通过观察和交互学习物理与因果,1000+异构游戏基准测试揭秘突破!
arXiv:2511.15407v3 Announce Type: replace Abstract: Humans learn by observing, interacting with environments, and internalizing physics and causality.…