1
MLLMs Get It Right, Then Get It Wrong: Tracing and Correcting Late-Layer Textual Bias
多模态大模型在图像矛盾时为何始终偏向文本?最新研究定位到晚层文本偏差并提出纠正方法
arXiv:2606.17953v1 Announce Type: new Abstract: When vision contradicts text, multimodal large language models (MLLMs) consistently favor text, even w…