Most of the LLM routing gap is task type
揭秘LLM路由差距的根源:任务类型比模型选择更关键,实测14模型两次运行结果并不完全一致。
arXiv:2608.23023v1 Announce Type: new Abstract: An LLM router picks which model should answer each query. The appeal is that models fail on different …
揭秘LLM路由差距的根源:任务类型比模型选择更关键,实测14模型两次运行结果并不完全一致。
arXiv:2608.23023v1 Announce Type: new Abstract: An LLM router picks which model should answer each query. The appeal is that models fail on different …
精准识别生物医学论文中的AI辅助写作痕迹,为学术诚信提供数据支撑,多引擎检测更可靠
IT之家 8 月 25 日消息,据外媒 phys 报道,近期 Lena Holzwarth 带头的研究团队分析了来自 PubMed Central 平台的 119 万余篇英文生物医学论文,发现生成式 AI 正迅速融入学术论文写作。 研究人员通过分析论文中的“AI 高频用词”(例“delves”“ex…
揭示低资源语言在向量空间中的几何结构,为多语言模型优化提供全新视角。
arXiv:2608.23358v1 Announce Type: new Abstract: The performance gap between low- and high-resource languages in LLMs is widely known, but it remains u…
针对多语言提示下大模型表现不一的难题,用定向合成数据提升模型智能,值得关注。
arXiv:2608.15964v1 Announce Type: cross Abstract: Language-specific competency (LSC) is the phenomenon of a language model performing better or worse …
一份系统文献综述,聚焦低资源语言下大模型安全对齐的挑战、方法与研究空白,适合关注多语言AI安全的研究者。
arXiv:2608.14626v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safet…
最大开源阿拉伯语大模型家族Jais 2发布,专精阿拉伯文化场景,性能对标国际主流模型。
arXiv:2608.13580v1 Announce Type: cross Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, an…
为印度亿万用户打造的多语言语音金融助手,看AI如何打破语言与信息壁垒
How I Built RupeeGPT: A Voice-First AI Financial Assistant for India in 10 Days #VoiceForBharat | Built with the fastest TTS API — Murf Falcon | 10 Da…
多语言环境下视觉语言模型的概念绑定稳定性大考,揭示跨语言推理的隐藏缺陷。
arXiv:2608.12333v1 Announce Type: cross Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associa…
5天从零搭建生产级Web工作室网站,AI聊天、多语言、分析全都有,全流程实战拆解
How I Built a Complete Web Studio Site in 5 Days Building a website is easy. Building a complete website with AI chat, multi-language support, real-ti…
越南语六大方言实测大模型鲁棒性,多语言能力短板一目了然。
arXiv:2608.10414v1 Announce Type: cross Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday co…
多语言预训练并非总是有益,这项研究揭示了LLM预训练不稳定的关键发现。
arXiv:2608.08800v1 Announce Type: new Abstract: Pretraining LLMs on artificial languages ("pre-pretraining") is a technique that could reportedly incr…
识别AI生成代码,支持多语言,帮开源项目守住新规红线。
IT之家 8 月 9 日消息,甲骨文(IT之家注:Oracle)现已通知 OpenJDK 开发者,要求项目组不能再提交 AI 生成的代码。 甲骨文对此表示:“ OpenJDK 社区贡献内容不得包含由大语言模型 、 扩散模型或深度学习系统生成的内容 。此处的‘内容’包括但不限于代码、文本、PR、电子邮…
新方法让大模型多语言数学推理更强,策略蒸馏带来显著提升,值得一看。
arXiv:2608.05802v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LL…
一条命令编排多语言monorepo构建,Rust编写,支持Elixir、TypeScript、Go,值得一试。
We built Aster to support our polyglot monorepo 'firstlanding'. At ArchAstro, we use Elixir, Python, TypeScript, Go, and Rust. I wanted something that…
一键将JS、Python、Rust等13种语言项目重写为Go代码,支持接入Claude Code、Codex等AI编程助手。
Article URL: https://github.com/JohnVictorCrown/WaterToGo Comments URL: https://news.ycombinator.com/item?id=49189271 Points: 1 # Comments: 2
IT之家 8 月 6 日消息,微软昨日(8 月 5 日)更新推出 Visual Studio Code 1.132 版本, 在内置浏览器中引入了元素级反馈、多语言语音输入、侧边聊天功能,并在混合 Markdown 编辑器中加入了 Markdown 差异比较功能。 在内置浏览器方面,VS …
多语言环境下多智能体规划为何频频“翻车”?这篇论文给出可操作的失效诊断,直击任务关键信息丢失的根源。
arXiv:2608.03735v1 Announce Type: cross Abstract: Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rare…
42项基准全面检验OpenAI隐私过滤器,跨语言跨域PII检测能力一探究竟,数据安全必读。
arXiv:2608.02616v1 Announce Type: new Abstract: We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-param…
让Claude接管视频剪辑,MCP协议驱动编辑器,自动配音字幕,30分钟出片,效率提升10倍
Article URL: https://www.shorz.ai Comments URL: https://news.ycombinator.com/item?id=49175711 Points: 2 # Comments: 0
用合成数据驯服农业领域多语言大模型,破解低资源语言问答难题,实验严谨、方法可复用。
arXiv:2507.16974v3 Announce Type: replace-cross Abstract: Enabling farmers to access accurate agriculture-related information in their native language…