RAP: KV-Cache Compression via RoPE-Aligned Pruning
直击长上下文推理痛点,用RoPE对齐剪枝实现KV-Cache压缩,兼顾效率与精度。
arXiv:2602.02599v4 Announce Type: replace-cross Abstract: Long-context inference in large language models (LLMs) is bottlenecked by the memory and com…
直击长上下文推理痛点,用RoPE对齐剪枝实现KV-Cache压缩,兼顾效率与精度。
arXiv:2602.02599v4 Announce Type: replace-cross Abstract: Long-context inference in large language models (LLMs) is bottlenecked by the memory and com…
固定KV缓存内存下,智能调度不同长度请求,突破LLM推理效率瓶颈。
arXiv:2508.06133v4 Announce Type: replace-cross Abstract: We study offline scheduling for large language model (LLM) serving under a fixed KV-cache me…
冷MoE模型多LLM服务新方案,KV-Cache与权重分离双管齐下,性能瓶颈有望突破。
arXiv:2606.24506v1 Announce Type: cross Abstract: Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse reque…
语义检索引导的KV缓存压缩,让长上下文LLM推理更省资源。
arXiv:2606.24467v1 Announce Type: new Abstract: Long-context large language model (LLM) inference is increasingly constrained by the memory footprint …
提出跨层KV-Cache压缩新方法xKV,利用对齐奇异向量提取大幅减少显存占用,加速大模型推理。
arXiv:2503.18893v2 Announce Type: replace-cross Abstract: Long-context Large Language Models (LLMs) enable powerful applications but incur high memory…