Benchmarking Patent Drafting from Inventor-Style Disclosures
科研新基准:从发明人式披露生成专利文本,直击大模型专业写作能力短板。
arXiv:2608.21249v1 Announce Type: new Abstract: While recent large language models (LLMs) have achieved promising results on individual patent draftin…
科研新基准:从发明人式披露生成专利文本,直击大模型专业写作能力短板。
arXiv:2608.21249v1 Announce Type: new Abstract: While recent large language models (LLMs) have achieved promising results on individual patent draftin…
前沿AI模型的政治立场、伦理取向与人格特质,一次基准测试看透底牌
Article URL: https://www.blackbench.ai/ Comments URL: https://news.ycombinator.com/item?id=49382540 Points: 3 # Comments: 2
安全专家实测122次AI攻防挑战,发现10次违规行为,揭示大模型在安全任务中的失控风险。
The AI Security Institute has a new report of AI systems engaging in “unsanctioned behavior”—what I have been calling “ genie …
让大模型在推理时像人一样从经验中迭代学习,Chain-of-Experience开启持续进化新范式
arXiv:2608.18027v1 Announce Type: new Abstract: Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations…
如何评估大模型在移动端处理零散个人信息的能力?这项研究给出了全新评测基准与洞见。
arXiv:2608.10692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is …
首个针对法律领域大模型时间推理能力的基准测试,填补评测空白
arXiv:2608.09106v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. …
大模型评测如何避免“自说自话”?一文拆解基准测试信任危机与去中心化验证路径。
arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are…
视频大模型缩放规律新发现:稳定曲线下藏着不稳定的样本差异,读懂这项研究才能避开评估陷阱。
arXiv:2608.07014v1 Announce Type: new Abstract: Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. …
提出“技能熵”概念,为长程推理的评估与训练提供新视角,值得关注。
arXiv:2608.05139v1 Announce Type: cross Abstract: Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a…
从玄学到科学:用户凭感觉“vibe-test”大模型,本文系统梳理并形式化这种评估方式,为LLM实用评测提供新视角。
arXiv:2604.14137v3 Announce Type: replace-cross Abstract: Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world…
LLM智能体能否真正持续进化?首个持续技能学习基准框架,直击智能体能力跃迁核心难题
arXiv:2608.03874v1 Announce Type: cross Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex t…
金融领域首个评估大模型工具调用能力的基准,看智能体能否玩转真实金融工具
arXiv:2603.08262v2 Announce Type: replace Abstract: The integration of Large Language Models (LLMs) into the financial domain is driving a paradigm sh…
聚焦金融、医疗、法律三大高风险领域,TRIDENT基准测试揭示LLM安全隐患与边界。
arXiv:2507.21134v2 Announce Type: replace-cross Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, …
从医学文本切入,揭示LLM水印技术的缺陷与挑战,直击当前AI安全盲区。
arXiv:2607.20462v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need f…
首个俄语用户记忆基准RUMBA,专门评测大模型在俄语场景下的记忆力表现,填补非英语评估空白。
arXiv:2607.21447v1 Announce Type: cross Abstract: The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing bench…
研究揭开了LLM在识别人类价值观时的困惑点,基于Schwartz理论的首个系统评估实验
arXiv:2607.20270v1 Announce Type: new Abstract: Large language models are increasingly evaluated through the values they endorse, but such evaluations…
首个针对吉尔吉斯语的大模型理解能力评测基准,填补低资源语言评估空白。
arXiv:2607.17173v1 Announce Type: new Abstract: Evaluating large language models (LLMs) across languages remains challenging, as most multilingual ben…
新基准Imaging-101系统评估LLM编码代理在科学计算成像上的表现,揭示当前模型能力边界。
arXiv:2607.10789v1 Announce Type: new Abstract: Computational imaging, which recovers hidden signals from indirect, noisy measurements, underpins quan…
健康问询里藏着文化暗号,CCBENCH 用隐式规范信号检验大模型能否不靠刻板印象读懂用户。
arXiv:2607.05405v1 Announce Type: cross Abstract: To interact with users fairly and without stereotyping, AI models must display cultural competency, …
ICML 2026论文为LLM Agent在社交困境中的合作能力打造首个系统化基准,揭示维持机制的关键设计。
arXiv:2604.15267v2 Announce Type: replace-cross Abstract: It is increasingly important that LLM agents interact effectively and safely with other goal…