Pallas: A Proactive KV Cache Migration Framework for LLM Inference in AI-RAN
首个面向AI-RAN的主动式KV缓存迁移框架,直击LLM推理中长序列与多边缘节点的时延痛点,架构设计值得关注。
arXiv:2608.16477v1 Announce Type: new Abstract: AI-RAN brings large language model (LLM) serving close to mobile users, but cellular handover can sepa…