The Geometry of Low-Resource Language Representations
揭示低资源语言在向量空间中的几何结构,为多语言模型优化提供全新视角。
arXiv:2608.23358v1 Announce Type: new Abstract: The performance gap between low- and high-resource languages in LLMs is widely known, but it remains u…
揭示低资源语言在向量空间中的几何结构,为多语言模型优化提供全新视角。
arXiv:2608.23358v1 Announce Type: new Abstract: The performance gap between low- and high-resource languages in LLMs is widely known, but it remains u…
一份系统文献综述,聚焦低资源语言下大模型安全对齐的挑战、方法与研究空白,适合关注多语言AI安全的研究者。
arXiv:2608.14626v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safet…
首个开源8B意第绪语大模型,并配套评测基准,填补稀缺语言建模空白。
arXiv:2608.05850v1 Announce Type: cross Abstract: We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yid…
揭示大模型在孟加拉语脏话上的安全漏洞:对齐仅绑定高资源形式,而非有害含义。
arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six prot…
用合成数据驯服农业领域多语言大模型,破解低资源语言问答难题,实验严谨、方法可复用。
arXiv:2507.16974v3 Announce Type: replace-cross Abstract: Enabling farmers to access accurate agriculture-related information in their native language…
首个针对吉尔吉斯语的大模型理解能力评测基准,填补低资源语言评估空白。
arXiv:2607.17173v1 Announce Type: new Abstract: Evaluating large language models (LLMs) across languages remains challenging, as most multilingual ben…
为低资源语言挑选实时语音AI技术栈的实战指南,直击冷门语种落地痛点
Article URL: https://kamalg2.substack.com/p/choosing-a-real-time-voice-ai-stack Comments URL: https://news.ycombinator.com/item?id=48814001 Points: 1 …
卢森堡语口语问答新突破,用TTS增强数据实现低资源语言语音交互。
arXiv:2607.02763v1 Announce Type: new Abstract: Spoken Question Answering (SQA) remains largely focused on high-resource languages and carefully recor…
面向低资源语言罗马尼亚语的多模态指令微调,用参数高效方法实现视觉语言模型适配,填补非英语VLM研究空白。
arXiv:2512.14926v2 Announce Type: replace-cross Abstract: Focusing on low-resource languages is an essential step toward democratizing generative AI. …
揭露低资源语言与混合代码切换可绕过LLM安全防护,揭示模型安全泛化的关键漏洞。
arXiv:2607.01859v1 Announce Type: new Abstract: Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncert…
卢森堡语大模型指令调优数据集LuxIT正式开源,独创单语种子数据高效构建方案,为低资源语言NLP突破奠定基础。
arXiv:2510.24434v3 Announce Type: replace Abstract: The effectiveness of instruction-tuned Large Language Models (LLMs) is often limited in low-resour…
首个阿拉伯语-俄语平行语料库与LLM基准,打破科研语言壁垒,助力可持续发展知识交流。
arXiv:2606.30943v1 Announce Type: new Abstract: Russian and Arabic are among the major languages of scientific communication. Language barriers impede…
为卢森堡语打造的高表现力语音合成语料库,填补低资源语言情感语音研究空白。
arXiv:2606.31947v1 Announce Type: new Abstract: State-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low…
多阶段LLM流水线如何兼顾翻译质量与原文版式?以马拉地语为案例,为低资源语言文档翻译提供新思路。
arXiv:2606.28796v1 Announce Type: cross Abstract: Government documents in India are predominantly issued in regional languages such as Marathi, creati…
专攻乌尔都语数学推理的8B模型,为低资源语言AI推理开辟新路径
arXiv:2606.25568v1 Announce Type: new Abstract: Recent LLMs demonstrate strong mathematical reasoning capabilities, but existing gains rely heavily on…
探究数据混合与模型架构对非洲语言持续预训练的影响,为低资源语言建模提供前沿实证与设计指南。
arXiv:2601.06395v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly multilingual, yet open models continue to underperfo…
打破罗马乌尔都语模型词汇瓶颈,几何保持扩展法让低资源语言AI更强大。
arXiv:2606.22478v1 Announce Type: new Abstract: Multilingual Language Models like mBERT are widely used for low-resource NLP, yet their adaptation to …
大模型在非洲低资源语言豪萨语和丰贝语上的翻译表现如何?新基准测试揭示失败模式与评估指标可靠性。
arXiv:2606.22269v1 Announce Type: cross Abstract: We investigate the translation quality of current large language models (LLMs) for English-to-Hausa …
不用人类测试集,用轮询合成数据评估法选出最佳LLM生成器,助力低资源语言模型训练
arXiv:2510.06143v2 Announce Type: replace Abstract: LLMs are powerful generators of synthetic data, which are used for training smaller, specific mode…
探索提示工程在低资源语言乌克兰语上的语法纠错极限,揭示大模型与提示策略的边界。
arXiv:2606.09334v1 Announce Type: new Abstract: Fine-tuned Large Language Models (LLMs) dominate in Ukrainian grammatical error correction (GEC), whil…