VLM- and LLM-Driven Multi-Agent System for PET Image Denoising
多智能体协同视觉语言模型,革新PET医学影像去噪新路径
arXiv:2608.13791v1 Announce Type: cross Abstract: Positron emission tomography (PET) imaging suffers from limited spatial resolution and low signal-to…
多智能体协同视觉语言模型,革新PET医学影像去噪新路径
arXiv:2608.13791v1 Announce Type: cross Abstract: Positron emission tomography (PET) imaging suffers from limited spatial resolution and low signal-to…
多语言环境下视觉语言模型的概念绑定稳定性大考,揭示跨语言推理的隐藏缺陷。
arXiv:2608.12333v1 Announce Type: cross Abstract: Vision-language models must associate visual entities with textual attributes. Whether these associa…
大模型竟对五条腿的狗视而不见?这项研究用溯因偏好学习破解提示不敏感难题。
arXiv:2510.09887v3 Announce Type: replace Abstract: Vision and language models frequently ignore semantically critical input edits, defaulting to pret…
视觉语言模型如何轻量化?这项研究提出免重训练的任务无关剪枝方案,一次剪枝即可适配下游任务。
arXiv:2608.06901v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal task…
多模态大模型能否预判危险?这份基准测试从运动场景到安全风险,评测AI的主动风险推断能力。
arXiv:2608.05560v1 Announce Type: cross Abstract: Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evalua…
谷歌Gemini新模型将AI能力注入实体机器人,理解环境并与人类交互,迈向物理世界智能化。
The latest version of Google DeepMind's AI model includes a significant jump into “physical AGI.” But plopping AI into the real world comes with risks…
无需3D训练数据,模块化视图感知让大模型直接推理3D问答,机器人感知新思路。
arXiv:2607.28442v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled new pos…
前沿视觉语言模型竟无图编造诊断,且编造内容与患者身份相关,揭示AI偏见风险。
arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do n…
用低成本探针方法预测昂贵训练效果,快速筛选3D CT视觉语言模型的最佳编码器和压缩方案。
arXiv:2607.22771v1 Announce Type: new Abstract: Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-com…
提出一种简单域泛化方法,显著增强现代视觉语言模型下像素级图像篡改检测的鲁棒性。
arXiv:2607.18230v1 Announce Type: cross Abstract: Modern vision-language models (VLMs) have significantly improved image generation and editing capabi…
面向低资源语言罗马尼亚语的多模态指令微调,用参数高效方法实现视觉语言模型适配,填补非英语VLM研究空白。
arXiv:2512.14926v2 Announce Type: replace-cross Abstract: Focusing on low-resource languages is an essential step toward democratizing generative AI. …
用AdaBoost集成文本提示,为视觉语言模型带来新提升,ECCVal2026 Spotlight论文值得关注。
arXiv:2607.00684v1 Announce Type: new Abstract: The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the t…
新数据集SpatialMosaic专攻部分可见场景,补全多视图VLM空间推理短板。
arXiv:2512.23365v4 Announce Type: replace Abstract: Recent progress in Multimodal Large Language Models (MLLMs) has enabled 3D scene understanding and…
机器人触觉与视觉语言结合的多模态数据集,为触觉泛化和具身智能模型训练提供新基准。
arXiv:2606.31694v1 Announce Type: cross Abstract: For robots manipulating open-world objects, tactile representations must generalize to unseen materi…
视觉语言模型在测试时也能通过缩放计算量提升性能,这篇论文揭示了新的缩放规律。
arXiv:2606.28864v1 Announce Type: new Abstract: Test-time scaling is a paradigm where large models use additional compute at inference to achieve bett…
VLMs会看图规划吗?CoSPlan用场景图增量更新,专治视觉决策中的顺序规划盲区。
arXiv:2512.10342v3 Announce Type: replace Abstract: Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remain…
CLIP的注意力机制有缺陷?Mamba架构对比学习新方案CLIMP来了,彻底抛弃Transformer,直击虚假关联与计算瓶颈
arXiv:2601.06891v2 Announce Type: replace Abstract: Contrastive Language-Image Pre-training (CLIP) relies on Vision Transformers whose attention mecha…
用视觉语言模型自动设计基于势能的奖励函数,加速强化学习探索
arXiv:2606.27180v1 Announce Type: cross Abstract: Sparse rewards are inherently challenging for reinforcement learning agents as they lack intermediat…
大型视觉语言模型重校准后生成更小模型,提升效率和精度,新技术论文。
arXiv:2506.15681v4 Announce Type: replace Abstract: Recent advancements in vision-language models (VLMs) have leveraged large language models (LLMs) t…
给视觉语言模型装上「耳朵」,用音频信号提升视觉理解能力,多模态融合新思路
arXiv:2606.23763v1 Announce Type: cross Abstract: Recent work typically assesses vision--language consistency using attention distributions of answer-…