FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
长上下文LLM推理提速新方案,块稀疏注意力优化预填充阶段,值得关注。
arXiv:2608.19758v1 Announce Type: new Abstract: Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity …