1
Allocation Before Ranking: Decoupled Token Compression for OmniLLMs
让全模态大模型告别冗余Token,解耦压缩新思路,先分配后排序,推理效率可期!
arXiv:2608.01665v1 Announce Type: new Abstract: Token compression in OmniLLMs is typically posed as a single saliency-ranking problem: score each mult…