1
HERALD: High-Throughput Block Diffusion LLM Serving via CPU-GPU Cooperative KV Cache Retrieval
CPU-GPU协同KV缓存检索,让块扩散LLM服务吞吐突破瓶颈,值得系统研究者细读。
arXiv:2606.21633v1 Announce Type: new Abstract: Diffusion LLMs (dLLMs) improve GPU utilization over autoregressive decoding by generating multiple tok…