1
The case for disaggregated LLM serving
把大模型推理拆成预填充和解码两段,用网络传KV缓存,破解长上下文延迟痛点,值得一读。
Article URL: https://blog.doubleword.ai/when-to-disaggregate Comments URL: https://news.ycombinator.com/item?id=49269470 Points: 4 # Comments: 0
把大模型推理拆成预填充和解码两段,用网络传KV缓存,破解长上下文延迟痛点,值得一读。
Article URL: https://blog.doubleword.ai/when-to-disaggregate Comments URL: https://news.ycombinator.com/item?id=49269470 Points: 4 # Comments: 0