1
Sangam: Efficiently Serving Diffusion LLMs with the AR Stack
扩散LLM推理效率新突破:Sangam用AR Stack化解双向注意力缓存难题,大幅加速生成服务。
arXiv:2607.04206v1 Announce Type: cross Abstract: Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can c…