VLM- and LLM-Driven Multi-Agent System for PET Image Denoising
多智能体协同视觉语言模型,革新PET医学影像去噪新路径
arXiv:2608.13791v1 Announce Type: cross Abstract: Positron emission tomography (PET) imaging suffers from limited spatial resolution and low signal-to…
多智能体协同视觉语言模型,革新PET医学影像去噪新路径
arXiv:2608.13791v1 Announce Type: cross Abstract: Positron emission tomography (PET) imaging suffers from limited spatial resolution and low signal-to…
多语言环境下视觉语言模型的概念绑定稳定性大考,揭示跨语言推理的隐藏缺陷。
arXiv:2608.12333v1 Announce Type: cross Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associa…
大模型竟对五条腿的狗视而不见?这项研究用溯因偏好学习破解提示不敏感难题。
arXiv:2510.09887v3 Announce Type: replace Abstract: Vision and language models frequently ignore semantically critical input edits, defaulting to pret…
视觉语言模型如何轻量化?这项研究提出免重训练的任务无关剪枝方案,一次剪枝即可适配下游任务。
arXiv:2608.06901v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal task…
用持久世界-自我状态建模破解VLA泛化难题,让机器人更懂环境与自身轨迹
arXiv:2608.06729v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive pa…
多模态大模型能否预判危险?这份基准测试从运动场景到安全风险,评测AI的主动风险推断能力。
arXiv:2608.05560v1 Announce Type: cross Abstract: Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evalua…
谷歌Gemini新模型将AI能力注入实体机器人,理解环境并与人类交互,迈向物理世界智能化。
The latest version of Google DeepMind's AI model includes a significant jump into “physical AGI.” But plopping AI into the real world comes with risks…
无需3D训练数据,模块化视图感知让大模型直接推理3D问答,机器人感知新思路。
arXiv:2607.28442v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled new pos…
前沿视觉语言模型竟无图编造诊断,且编造内容与患者身份相关,揭示AI偏见风险。
arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do n…
用低成本探针方法预测昂贵训练效果,快速筛选3D CT视觉语言模型的最佳编码器和压缩方案。
arXiv:2607.22771v1 Announce Type: new Abstract: Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-com…
提出一种简单域泛化方法,显著增强现代视觉语言模型下像素级图像篡改检测的鲁棒性。
arXiv:2607.18230v1 Announce Type: cross Abstract: Modern vision-language models (VLMs) have significantly improved image generation and editing capabi…
多模态大模型在复杂交通场景中常忽略关键小目标,这项研究提出新基准与查询引导聚焦机制,让AI真正看见细节。
arXiv:2607.04149v1 Announce Type: new Abstract: In safety-critical traffic scenarios, answering complex questions relies on minute, localized visual c…
用文本精准描述3D场景,为多模态理解开辟新方向,论文含金量十足。
arXiv:2607.02908v1 Announce Type: new Abstract: This work introduces holo-captioning, a novel task that strives to seek the text equivalent of 3D scen…
面向低资源语言罗马尼亚语的多模态指令微调,用参数高效方法实现视觉语言模型适配,填补非英语VLM研究空白。
arXiv:2512.14926v2 Announce Type: replace-cross Abstract: Focusing on low-resource languages is an essential step toward democratizing generative AI. …
用AdaBoost集成文本提示,为视觉语言模型带来新提升,ECCVal2026 Spotlight论文值得关注。
arXiv:2607.00684v1 Announce Type: new Abstract: The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the t…
新数据集SpatialMosaic专攻部分可见场景,补全多视图VLM空间推理短板。
arXiv:2512.23365v4 Announce Type: replace Abstract: Recent progress in Multimodal Large Language Models (MLLMs) has enabled 3D scene understanding and…
机器人触觉与视觉语言结合的多模态数据集,为触觉泛化和具身智能模型训练提供新基准。
arXiv:2606.31694v1 Announce Type: cross Abstract: For robots manipulating open-world objects, tactile representations must generalize to unseen materi…
视觉语言模型在测试时也能通过缩放计算量提升性能,这篇论文揭示了新的缩放规律。
arXiv:2606.28864v1 Announce Type: new Abstract: Test-time scaling is a paradigm where large models use additional compute at inference to achieve bett…
掩码引导区域感知预训练,让AI更精准读懂胎儿超声影像,医学视觉语言模型新突破。
arXiv:2606.29586v1 Announce Type: new Abstract: Vision-language foundation models have shown strong potential in medical image analysis. Although foun…
VLMs会看图规划吗?CoSPlan用场景图增量更新,专治视觉决策中的顺序规划盲区。
arXiv:2512.10342v3 Announce Type: replace Abstract: Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remain…