Most of the LLM routing gap is task type
揭秘LLM路由差距的根源:任务类型比模型选择更关键,实测14模型两次运行结果并不完全一致。
arXiv:2608.23023v1 Announce Type: new Abstract: An LLM router picks which model should answer each query. The appeal is that models fail on different …
揭秘LLM路由差距的根源:任务类型比模型选择更关键,实测14模型两次运行结果并不完全一致。
arXiv:2608.23023v1 Announce Type: new Abstract: An LLM router picks which model should answer each query. The appeal is that models fail on different …
科研新基准:从发明人式披露生成专利文本,直击大模型专业写作能力短板。
arXiv:2608.21249v1 Announce Type: new Abstract: While recent large language models (LLMs) have achieved promising results on individual patent draftin…
前沿AI模型的政治立场、伦理取向与人格特质,一次基准测试看透底牌
Article URL: https://www.blackbench.ai/ Comments URL: https://news.ycombinator.com/item?id=49382540 Points: 3 # Comments: 2
安全专家实测122次AI攻防挑战,发现10次违规行为,揭示大模型在安全任务中的失控风险。
The AI Security Institute has a new report of AI systems engaging in “unsanctioned behavior”—what I have been calling “ genie …
让大模型在推理时像人一样从经验中迭代学习,Chain-of-Experience开启持续进化新范式
arXiv:2608.18027v1 Announce Type: new Abstract: Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations…
深度学习中模型学到的是真实生物信号还是伪影?ICML 2026 论文以生物成像为例,拷问表征学习的本质,值得每一位科研工作者反思。
arXiv:2603.13377v2 Announce Type: replace-cross Abstract: Representation learning has driven major advances in natural image analysis by enabling mode…
模型升级未必全赢:新研究揭示样本级回归无法用通用信号预测,AI迭代背后的隐忧值得关注。
arXiv:2608.13607v1 Announce Type: new Abstract: Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate. But agg…
探究大模型能否识别自身易错场景,为AI安全与可靠部署提供新视角。
arXiv:2607.23496v2 Announce Type: replace Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the sam…
更多次搜索竟比更强搜索模型更有效?OpenRouter基准测试揭示LLM搜索策略的真相。
Article URL: https://openrouter.ai/blog/announcements/web-search-benchmark/ Comments URL: https://news.ycombinator.com/item?id=49302033 Points: 2 # Co…
不用等生成完,从预生成激活就能预判模型成败,给LLM装上“事前质检仪”。
arXiv:2602.09924v4 Announce Type: replace-cross Abstract: Running LLMs with extended reasoning on every problem is expensive, but determining which in…
如何评估大模型在移动端处理零散个人信息的能力?这项研究给出了全新评测基准与洞见。
arXiv:2608.10692v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is …
LLM推理并非万能,主观任务中会失灵?这篇论文剖析失败模式并给出动态路由方案。
arXiv:2608.08889v1 Announce Type: new Abstract: Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a…
首个针对法律领域大模型时间推理能力的基准测试,填补评测空白
arXiv:2608.09106v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. …
大模型评测如何避免“自说自话”?一文拆解基准测试信任危机与去中心化验证路径。
arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are…
首个用“人生事件”测试LLM人格演化的基准,揭示AI人格能否像人一样经历成长与改变。
arXiv:2608.06485v1 Announce Type: cross Abstract: Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social si…
视频大模型缩放规律新发现:稳定曲线下藏着不稳定的样本差异,读懂这项研究才能避开评估陷阱。
arXiv:2608.07014v1 Announce Type: new Abstract: Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. …
提出“技能熵”概念,为长程推理的评估与训练提供新视角,值得关注。
arXiv:2608.05139v1 Announce Type: cross Abstract: Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a…
从玄学到科学:用户凭感觉“vibe-test”大模型,本文系统梳理并形式化这种评估方式,为LLM实用评测提供新视角。
arXiv:2604.14137v3 Announce Type: replace-cross Abstract: Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world…
LLM智能体能否真正持续进化?首个持续技能学习基准框架,直击智能体能力跃迁核心难题
arXiv:2608.03874v1 Announce Type: cross Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex t…
42项基准全面检验OpenAI隐私过滤器,跨语言跨域PII检测能力一探究竟,数据安全必读。
arXiv:2608.02616v1 Announce Type: new Abstract: We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-param…