1
CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation
冷MoE模型多LLM服务新方案,KV-Cache与权重分离双管齐下,性能瓶颈有望突破。
arXiv:2606.24506v1 Announce Type: cross Abstract: Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse reque…