An AI4AI Framework for Visual Token Pruning
把Token剪枝交给AI自动设计,告别手工调参,多模态大模型推理成本有望再降一截。
arXiv:2608.07193v1 Announce Type: new Abstract: Visual-token pruning can substantially reduce the inference cost of multimodal large language models (…
把Token剪枝交给AI自动设计,告别手工调参,多模态大模型推理成本有望再降一截。
arXiv:2608.07193v1 Announce Type: new Abstract: Visual-token pruning can substantially reduce the inference cost of multimodal large language models (…
多模态大模型提速新思路:按角色分区裁剪视觉token,兼顾效率与精度
arXiv:2608.07088v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefil…
多模态大模型有惊喜:砍掉一半视觉token反而更鲁棒?最新研究揭秘token压缩的抗干扰机制。
arXiv:2607.22716v1 Announce Type: cross Abstract: In this paper, we show for the first time that visual token pruning enhances the robustness of Multi…
多模态大模型新方法,实现视觉token高效嵌入,已被ECCV 2026录用。
arXiv:2602.05275v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have shown immense promise in universal multimodal retrie…
熵引导视觉token剪枝,修正注意力机制,让多模态大模型推理更快更省。
arXiv:2606.31982v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) incur prohibitive inference costs due to long visual token se…
用第一性原理构建最优保留集,为多模态大模型推理大幅削减视觉token,效率提升新思路。
arXiv:2606.27161v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but t…
多模态大模型推理成本主要来自视觉token,提出逐步选择方法大幅降低计算量,实现高效生成。
arXiv:2606.16067v1 Announce Type: new Abstract: In multimodal large language models (MLLMs), inference cost is largely dominated by the visual token p…
提出将不确定性引入动态时间规整,同时处理时序序列与视觉Token对齐,提升鲁棒性。
arXiv:2605.25110v1 Announce Type: cross Abstract: Aligning structured data is a fundamental problem in computer vision and machine learning, underlyin…
无需额外训练,通过空间-时间池化与网格化巧妙提升视频大语言模型视觉token表征,ICLR 2026接收!
arXiv:2605.22078v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have significantly advanced video unders…