Efficient Clustering with Provable Guardrails for LLM Inference at Scale
提出可证明安全的高效聚类方案,为大规模LLM推理保驾护航。
arXiv:2607.19704v1 Announce Type: new Abstract: Scaling LLM-based applications to millions of users is bottlenecked by the inference cost and latency …