PixVL: Self-Supervised Training of Pixel-Level MLLMs via a Unified Mask--Text Consistency Cycle
像素级多模态大模型新训练范式,掩码—文本一致性循环让模型精修细节,值得细读。
arXiv:2608.01354v1 Announce Type: new Abstract: Recent studies develop pixel-level multimodal large language models (MLLMs) that support both Region S…