1
VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs
无需训练即可实现可变率KV缓存压缩,显著降低长上下文LLM推理内存开销,附大量实验验证效果。
arXiv:2607.15498v1 Announce Type: cross Abstract: The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) in…