1
Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
千亿级大模型推理下沉到PC集群,预编译管道分片让闲置AI PC变身分布式算力池,技术方案直击成本痛点。
arXiv:2608.19147v1 Announce Type: cross Abstract: Modern Intel AI PCs ship capable integrated GPUs and NPUs with 16+ GB of unified memory, and they sp…