Remote-Sensing City Layout Extraction with MLLM
用多模态大模型把遥感影像直接变成可执行的程序化城市布局,打通识别到重建的闭环。
arXiv:2608.16484v1 Announce Type: new Abstract: Remote-sensing systems usually describe urban content with detection boxes, semantic masks, or vector …
用多模态大模型把遥感影像直接变成可执行的程序化城市布局,打通识别到重建的闭环。
arXiv:2608.16484v1 Announce Type: new Abstract: Remote-sensing systems usually describe urban content with detection boxes, semantic masks, or vector …
跨房间3D场景理解遇上拓扑感知多模态大模型,突破单一空间局限,推动空间智能新边界。
arXiv:2607.06534v1 Announce Type: new Abstract: Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in si…
3D场景图配准高斯溅射,用语义关联提升点云对齐精度,图形学新突破。
arXiv:2606.29782v1 Announce Type: new Abstract: Merging multiple 3D Gaussian Splatting (3DGS) scenes into a single unified Gaussian representation is …
异构视角下检视多模态大模型空间智能,这份基准或成具身协作关键标尺。
arXiv:2606.28049v1 Announce Type: new Abstract: In recent years, multimodal large language models (MLLMs) have shown strong potential for embodied int…
用AirPods和3D空间智能实时镜像运动姿态的开源体态教练,小熊化身陪你跑跳拉伸
Article URL: https://airposture.github.io Comments URL: https://news.ycombinator.com/item?id=48673814 Points: 2 # Comments: 0
清华开源Spatial-TTT,让模型在视频中持续学习更新空间记忆,空间智能超越Gemini,入选ECCV 2026!
120分钟长视频一边看一边记
训练多模态大模型直接内化3D空间感知,用可解释监督替代外部工具,降本增效,值得关注。
arXiv:2606.19915v1 Announce Type: new Abstract: Unlocking the spatial intelligence of multimodal large language model (MLLMs) is crucial for understan…
原生3D空间理解与生成模型,输出可编辑的3DGS格式,直接导入Unity/Unreal引擎,将AI转化为游戏开发生产力。
能够直接导入Unity、Unreal Engine等主流引擎进行交互开发
新框架CoCoSI让多个智能体协作构建认知地图,突破空间智能瓶颈
arXiv:2606.10401v1 Announce Type: new Abstract: Spatial intelligence is a key frontier for multimodal large language models (MLLMs), enabling them to …
5秒搞定多视角一致的3D场景编辑,北大港中文联合提出VGGT-Edit,加速120倍,解决空间智能致命缺陷。
不再绕回2D
双路径几何感知机制提升MLLM空间理解,为具身智能与3D视觉任务提供新范式。
arXiv:2605.25334v1 Announce Type: new Abstract: Spatial understanding of the physical world from 2D visual inputs hinges on two complementary forms of…
首个专注多图像空间智能的VQA基准,填补MLLM多图关系评估空白
arXiv:2505.23764v3 Announce Type: replace-cross Abstract: Spatial intelligence is essential for multimodal large language models (MLLMs) operating in …
一个专门用来评测具身空间智能的新基准
统一多模态理解与生成框架,唤醒空间智能,为AI视觉注入全新维度。
arXiv:2605.04128v2 Announce Type: replace-cross Abstract: We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text…
用数据集+模型+基准全方位提升多模态大模型跨视图空间智能,突破单视角局限。
arXiv:2605.18621v1 Announce Type: new Abstract: Spatial intelligence requires multimodal large language models (MLLMs) to move beyond single-view perc…
多模态大模型在空间智能上的突破,赋予AI更强的视觉感知与推理能力。
arXiv:2505.23747v2 Announce Type: replace-cross Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced …