1
LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models
新方法LOCUS通过局部视觉线索搜索,大幅提升多模态大模型对图像细节的感知能力
arXiv:2606.16586v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even whe…