Audio-Based Understanding of Audiobook Narration Appeal
用音频信号解码有声书旁白的吸引力,为配音选角与内容推荐提供新量化维度。
arXiv:2607.02473v1 Announce Type: new Abstract: Narration is central to the audiobook listening experience, shaping how listeners engage with and unde…
用音频信号解码有声书旁白的吸引力,为配音选角与内容推荐提供新量化维度。
arXiv:2607.02473v1 Announce Type: new Abstract: Narration is central to the audiobook listening experience, shaping how listeners engage with and unde…
视频大语言模型是否真的需要音频?这篇被Interspeech 2026接收的研究给出了基准审计与可扩展修复方案。
arXiv:2509.17901v4 Announce Type: replace Abstract: Speech and audio encoders developed over years of community effort are routinely excluded from vid…
新研究揭秘大型音频语言模型是否忠实于输入音频,为模型可信度提供关键评估。
arXiv:2509.22363v4 Announce Type: replace Abstract: Large Audio Language Models (LALMs) integrate audio encoders with pretrained Large Language Models…
首次将空间音频理解融入多模态大语言模型,通过FOA编码实现突破性整合。
arXiv:2606.10738v1 Announce Type: cross Abstract: Recent multimodal large language models mainly process audio as monaural signals, thereby discarding…
用LoRA技术让大模型直接理解音频,高效内化跨模态能力,开源研究新突破。
arXiv:2606.11033v1 Announce Type: cross Abstract: Recent efforts to extend large language models (LLMs) to speech inputs typically rely on cascaded AS…
MOSS大模型音频能力技术报告,揭秘音频理解与生成新突破
arXiv:2606.01802v1 Announce Type: cross Abstract: MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understandin…
研究发现视频MLLMs的音频理解实际依赖视觉线索,揭示模型幻觉问题,挑战多模态真实性。
arXiv:2605.16403v1 Announce Type: new Abstract: Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in vide…