豆包输入法鸿蒙版 0.9.1 上架尝鲜,语音输入可精准识别快速 / 轻声说话
豆包输入法鸿蒙版尝鲜,语音大模型加持,快慢轻声都能精准识别,还支持无网使用!
IT之家 8 月 22 日消息,豆包输入法鸿蒙版现已在 App Gallery 应用商店 开启尝鲜 ,版本号为 0.9.1(901),时间为 8 月 22 日至 9 月 21 日。 官方表示,豆包输入法的「语音输入」使用豆包同款语音大模型,语音输入又快又准,支持多方言、英语及中英混输;快速说话、轻声…
豆包输入法鸿蒙版尝鲜,语音大模型加持,快慢轻声都能精准识别,还支持无网使用!
IT之家 8 月 22 日消息,豆包输入法鸿蒙版现已在 App Gallery 应用商店 开启尝鲜 ,版本号为 0.9.1(901),时间为 8 月 22 日至 9 月 21 日。 官方表示,豆包输入法的「语音输入」使用豆包同款语音大模型,语音输入又快又准,支持多方言、英语及中英混输;快速说话、轻声…
低资源场景下语音大模型的数据需求,以及高资源语言预训练的影响,实验数据扎实,值得一读。
arXiv:2508.05149v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-…
FBK提出的长篇语音大模型在IWSLT 2026指令跟随任务上的探索,展示如何用LLM处理长语音输入。
arXiv:2606.26819v1 Announce Type: new Abstract: This paper describes our submission to the IWSLT 2026 Instruction Following shared task. SpeechLLMs ar…
探讨翻译增强预训练对语音大模型的实际影响,为语音AI研究提供实证依据。
arXiv:2606.25444v1 Announce Type: cross Abstract: Connecting a pre-trained speech encoder to a Large Language Model (LLM) is the standard architecture…
最新研究量化了语音LLM中的交叉偏见,揭示声音特征与种族、性别等多重因素的交互影响,为AI公平性提供关键评估方法。
arXiv:2603.16941v2 Announce Type: replace-cross Abstract: Speech Large Language Models (SpeechLLMs) process spoken input directly, retaining cues such…
端到端语音LLM性能退化难题,这篇论文提出跨模态蒸馏方法对齐能力,直击行业痛点。
arXiv:2603.24596v3 Announce Type: replace-cross Abstract: While the shift from cascaded dialogue systems to end-to-end (E2E) speech Large Language Mod…
探讨文本作为信息瓶颈在语音大模型中的核心作用,为连续声学信号集成提供新思路
arXiv:2606.09366v1 Announce Type: new Abstract: Large language models (LLMs) provide a powerful reasoning backbone for speech understanding, but integ…
微调语音大模型,实现多粒度第二语言评估与自然语言理由生成,教育AI新突破。
arXiv:2606.09470v1 Announce Type: new Abstract: Automated L2 speech assessment can assign proficiency labels, but often lacks interpretability. We pro…
语音大模型推理中实体绑定出错了怎么办?这篇论文系统诊断问题并提出思维链干预方案。
arXiv:2606.04474v1 Announce Type: new Abstract: Speech Large Language Models (SLLMs) underperform their text counterparts on complex reasoning. We rev…
无需训练的纯解码器注意力新策略,让语音大模型长文本同声翻译更高效稳定。
arXiv:2605.31432v1 Announce Type: cross Abstract: Simultaneous speech-to-text translation (SimulST) generates translations while speech is still unfol…