Most of the LLM routing gap is task type
揭秘LLM路由差距的根源:任务类型比模型选择更关键,实测14模型两次运行结果并不完全一致。
arXiv:2608.23023v1 Announce Type: new Abstract: An LLM router picks which model should answer each query. The appeal is that models fail on different …
揭秘LLM路由差距的根源:任务类型比模型选择更关键,实测14模型两次运行结果并不完全一致。
arXiv:2608.23023v1 Announce Type: new Abstract: An LLM router picks which model should answer each query. The appeal is that models fail on different …
一句话看清大模型“变脸”真相:提示词微调竟让输出天翻地覆,交互式评估方法论带你量化并解释这种敏感脆弱。
arXiv:2608.18539v1 Announce Type: cross Abstract: The remarkable capabilities of large language models (LLMs) are often undermined by their instabilit…
Claude 又现529错误,服务稳定性引热议,关注大模型运维挑战
Comments URL: https://news.ycombinator.com/item?id=49362857 Points: 2 # Comments: 1
钉住GitHub服务波动数据,了解七月八个影响可用性的事件细节。
In July, we experienced eight incidents that resulted in degraded performance across GitHub services. The post GitHub availability report: July 2026 a…
多语言预训练并非总是有益,这项研究揭示了LLM预训练不稳定的关键发现。
arXiv:2608.08800v1 Announce Type: new Abstract: Pretraining LLMs on artificial languages ("pre-pretraining") is a technique that could reportedly incr…
衡量大模型微调后功能韧性的新指标,为数据适应期的稳定性评估提供量化工具。
arXiv:2608.03887v1 Announce Type: new Abstract: Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S…
OpenClaw 新版本修复了 Codex 进度回复中断与内存启动冲突,稳定性关键更新。
Fixes Codex progress replies: keep app-server turns running after delivered progress messages so GPT/Codex reaches its authoritative terminal response…
用 Laravel 实现熔断器,防止级联故障,保障微服务稳定性。
The Microservice Domino Effect Modern enterprise applications rarely operate in isolation. Your Laravel backend is likely communicating with a myriad …
真实用户吐槽Claude Code频繁掉线,开发者需知的最前线稳定性报告
I have just had my 3rd session in less than a week when I had to eventually quit Claude Code and do a task using another model/subscription. Am I just…
探索Titan瞬态现象与LLM可扩展性的交叉前沿,揭开大规模系统性能瓶颈的核心见解。
Article URL: https://queue.acm.org/detail.cfm?id=3819082 Comments URL: https://news.ycombinator.com/item?id=49095428 Points: 2 # Comments: 0
最新研究提出神经网络训练中的自动稳定性与恢复方法,为模型训练失败提供智能解决方案。
arXiv:2601.17483v2 Announce Type: replace-cross Abstract: Training modern neural networks is increasingly fragile, with rare but severe destabilizing …
新基准StabilityBench系统评估LLM输出的不稳定性,揭示模型可靠性的关键挑战。
arXiv:2607.20558v1 Announce Type: cross Abstract: AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government se…
前沿大模型存在“响应漂移”?这篇论文系统研究了不同版本LLM输出随时间一致性的关键问题。
arXiv:2607.20454v1 Announce Type: cross Abstract: All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate f…
Koopman理论赋能世界模型,谱约束稳定想象生成,突破模型长期预测的不稳定瓶颈。
arXiv:2607.19719v1 Announce Type: new Abstract: Latent world models improve sample efficiency in continuous control by optimizing policies over imagin…
最新显卡驱动修复多款游戏问题,NVIDIA和英特尔齐出手,稳定畅玩游戏必备更新。
IT之家 7 月 23 日消息,NVIDIA(英伟达)当地时间 22 日发布了基于 Game Ready Driver 610.74 的 610.82 版本 GeForce 热修复显示驱动程序。 这一版本 解决了两款游戏中的问题 :GeForce RTX 50 系列显卡在使用此前的 610.xx 驱…
提出归一化奖励方法,提升偏好优化训练稳定性与效果
arXiv:2607.16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs…
低秩预训练大模型遭遇不稳定困境?这项ICML 2026研究提出原生低秩LLM预训练的稳定化方法,兼顾效率与质量。
arXiv:2602.12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose signif…
36氪获悉,苏州铂氢科技源科技有限公司(以下简称「铂氢科技」)宣布完成数千万元Pre-A轮融资。本轮融资由东运创投领投、新晖资本跟投,老股东熔拓资本追投,资金将主要用于催化剂与膜电极新产线建设、研发及运营团队扩充。 铂氢科技成立于2023年5月,专注于氢能等新兴领域的贵金属催化剂及膜电极等下游产品研…
提出“思维策略”框架,通过在线策略进化实现测试时训练,打破冻结策略限制,显著提升大模型复杂推理能力。
arXiv:2601.20379v2 Announce Type: replace Abstract: Large language models (LLMs) struggle with complex, long-horizon reasoning due to instability caus…
IT之家 7 月 14 日消息,科技媒体 Cult of Mac 今天(7 月 14 日)发布博文,报道称苹果最新发布的 iOS 27 首个公测版已基本完善,其稳定性足以满足日常使用需求。 该媒体测试后认为 iOS 27 首个公测版已趋于完善,各方面的体验都非常稳定,完全满足日常使用需求,不过测试后…