AI Security Policy Should Assess Systems, Not Only Models
多智能体协同攻击框架揭示:评估AI安全不能只看模型,更要看系统。
arXiv:2605.09504v2 Announce Type: replace-cross Abstract: We present swarm-attack, an open-source adversarial testing framework in which multiple ligh…
多智能体协同攻击框架揭示:评估AI安全不能只看模型,更要看系统。
arXiv:2605.09504v2 Announce Type: replace-cross Abstract: We present swarm-attack, an open-source adversarial testing framework in which multiple ligh…
系统评估LLM作为裁判时的偏见缓解策略,揭示不同方法的有效性,为构建公平AI评估体系提供关键指南。
arXiv:2604.23178v2 Announce Type: replace Abstract: LLM-as-a-Judge has become the dominant paradigm for evaluating language model outputs, yet LLM jud…
评估四大主流实时语音AI在“言外之意”上的表现,揭示模型“听得到词却听不懂调”的盲点
arXiv:2606.26083v1 Announce Type: new Abstract: Speech conveys information through both words and vocal delivery. We evaluate four leading production …
系统性对比多种大模型黑盒不确定性估计方法,助你理解LLM可靠性与风险量化。
arXiv:2606.19868v1 Announce Type: new Abstract: Although large language models (LLMs) have shown strong capabilities across a wide range of tasks, the…
探索AI编码代理能否系统性复现社会科学发现,最新研究填补评估空白。
arXiv:2606.11447v1 Announce Type: new Abstract: Recent anecdotal evidence suggests that AI coding agents can reproduce published findings when provide…
被ICML 2026接收的系统性研究,用科学方法建立AI Agent可靠性评估框架,附交互式仪表盘。
arXiv:2602.16666v3 Announce Type: replace Abstract: AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on st…
系统评估揭示LLM在饮食失调咨询中产生“食物噪音”和“虚假安全”问题,临床反馈仍不足以解决漏洞。
arXiv:2606.02444v1 Announce Type: new Abstract: Recent evidence shows that people with eating disorders (EDs) are increasingly seeking guidance, advic…
AI评估新范式:STABLEVAL在评估中引入分歧感知与稳定性,为AI系统的可靠评价提供创新方案。
arXiv:2605.02122v2 Announce Type: replace Abstract: Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disag…
Science杂志深度评测名为“Co-Scientist”的新AI科研系统,带你一窥AI如何变革科学发现流程。
Article URL: https://www.science.org/content/blog-post/evaluating-co-scientist-new-ai-science-system Comments URL: https://news.ycombinator.com/item?i…