1
A Spatio-Temporal Expert Prefetching Framework for Efficient MoE-based LLM Inference
突破性提出时空专家预取框架,显著提升MoE大模型推理效率,值得关注。
arXiv:2606.15453v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) based large language models (LLMs), such as Qwen and DeepSeek, have recentl…