1
Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference
多模态大模型推理太慢?这项研究从算子级入手,实现视觉信息智能跳过,兼顾效率与精度。
arXiv:2606.31903v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) increasingly process long visual-token sequences, increasin…