Remote-Sensing City Layout Extraction with MLLM
用多模态大模型把遥感影像直接变成可执行的程序化城市布局,打通识别到重建的闭环。
arXiv:2608.16484v1 Announce Type: new Abstract: Remote-sensing systems usually describe urban content with detection boxes, semantic masks, or vector …
用多模态大模型把遥感影像直接变成可执行的程序化城市布局,打通识别到重建的闭环。
arXiv:2608.16484v1 Announce Type: new Abstract: Remote-sensing systems usually describe urban content with detection boxes, semantic masks, or vector …
原生多模态成K3与DeepSeek V4分水岭,看懂AI模型演进新方向。
文 | 李炤锋 编辑 | 张雨忻 “长链任务如果只通过代码层面的反馈,误差可能会不断累积,最终效果会非常差。”谈及原生多模态的意义,一位多模态研究员表示,“视觉是一种更准确的反馈,也更贴近用户意图。” 过去一年,Coding与Agent能力不断改写大模型的排名,也成为AI最快兑现商业价…
多模态大模型时代,显著目标检测迎来复兴契机,这篇论文带你重新审视经典任务的未来方向。
arXiv:2607.29222v1 Announce Type: new Abstract: The zero-shot capabilities of multimodal large language models (MLLMs) are pushing salient object dete…
多模态大模型也能「压缩记忆」?这篇论文提出在不牺牲理解能力的前提下,高效压缩视觉与文本token,让全模态LLM更轻更快。
arXiv:2607.21179v1 Announce Type: new Abstract: The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLM…
商汤开源统一视觉大模型 SenseNova-Vision,全任务能力超越 Vision Banana,视觉理解与生成实现原生统一。
IT之家 7 月 13 日消息,商汤科技今日发布并全面开源 日日新 SenseNova-Vision 理解生成统一视觉大模型 ,这是商汤日日新大模型体系的重要视觉能力升级。 商汤科技表示,行业以往的 " 统一视觉 " 多是把检测、分割、深度预测等多个专家模型打包封装,本质还是割裂的。SenseNov…
用概念引导提升上下文分割的鲁棒性,为AI视觉理解提供更稳的新思路
arXiv:2606.28149v1 Announce Type: cross Abstract: In-context segmentation (ICS) requires a model to segment target regions in a query image using only…
大模型如何给视觉模型当老师?ICML 2026探讨细粒度知识跨模态迁移,让AI理解更精准。
arXiv:2606.27527v1 Announce Type: cross Abstract: Large Language Models (LLMs) possess broad conceptual knowledge acquired through large-scale text pr…
新基准DiCoBench专攻多图像细粒度感知,用差异与共性视觉线索考验模型隐含理解力。
arXiv:2606.26602v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive fine-grai…
MLLM主动感知纠错新方法:让模型像人类一样主动发现并修正视觉错误,提升多模态理解准确性。
arXiv:2606.24292v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated impressive vision-language understanding, y…
给视觉语言模型装上「耳朵」,用音频信号提升视觉理解能力,多模态融合新思路
arXiv:2606.23763v1 Announce Type: cross Abstract: Recent work typically assesses vision--language consistency using attention distributions of answer-…
给多模态模型做压力测试的歧义评估基准,帮研究者精准定位视觉语言理解的结构性盲区
arXiv:2606.19552v1 Announce Type: new Abstract: Structural ambiguity arises when a single sentence admits multiple valid interpretations due to its sy…
多模态大模型新框架:先候选发现再比较推理,解决复杂查询下的图像分割难题
arXiv:2606.09303v1 Announce Type: new Abstract: The rapid development of pretrained foundation models has enabled more general image segmentation. Mul…
论文提出ValueGround基准,评估多模态大模型对不同文化背景下的视觉价值理解能力,揭示现有模型在文化适应性上的不足。
arXiv:2604.06484v3 Announce Type: replace Abstract: Cultural values are expressed not only through language but also through visual scenes and everyda…
重新思考高分辨率多模态大模型中的Zoom-IN方法,提出Hierarchical Decoupling框架,显著提升视觉理解性能。
arXiv:2510.00054v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding tas…
多模态大模型在空间智能上的突破,赋予AI更强的视觉感知与推理能力。
arXiv:2505.23747v2 Announce Type: replace-cross Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced …