An AI4AI Framework for Visual Token Pruning
把Token剪枝交给AI自动设计,告别手工调参,多模态大模型推理成本有望再降一截。
arXiv:2608.07193v1 Announce Type: new Abstract: Visual-token pruning can substantially reduce the inference cost of multimodal large language models (…
RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs
多模态大模型提速新思路:按角色分区裁剪视觉token,兼顾效率与精度
arXiv:2608.07088v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefil…
Visual Token Compression Enhances Robustness of MLLMs
多模态大模型有惊喜:砍掉一半视觉token反而更鲁棒?最新研究揭秘token压缩的抗干扰机制。
arXiv:2607.22716v1 Announce Type: cross Abstract: In this paper, we show for the first time that visual token pruning enhances the robustness of Multi…
ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs
熵引导视觉token剪枝,修正注意力机制,让多模态大模型推理更快更省。
arXiv:2606.31982v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) incur prohibitive inference costs due to long visual token se…
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
用第一性原理构建最优保留集,为多模态大模型推理大幅削减视觉token,效率提升新思路。
arXiv:2606.27161v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but t…