1
Unified Static-Dynamic Pruning for Efficient LLM Inference
提出统一静态与动态剪枝的框架,显著提升大模型推理速度与资源效率。
arXiv:2607.21985v1 Announce Type: cross Abstract: The increasing deployment of large language models (LLMs) has magnified the computational and memory…