1
RAP: KV-Cache Compression via RoPE-Aligned Pruning
直击长上下文推理痛点,用RoPE对齐剪枝实现KV-Cache压缩,兼顾效率与精度。
arXiv:2602.02599v4 Announce Type: replace-cross Abstract: Long-context inference in large language models (LLMs) is bottlenecked by the memory and com…