1
WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs
针对长语音大模型KV缓存膨胀问题,提出WnW动态增减机制,让长语音推理更省显存、更流畅。
arXiv:2608.22704v1 Announce Type: new Abstract: Long-form audio inputs make the KV cache the dominant memory cost of speech LLMs. Prefill-only KV comp…