1
Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs
无需训练即可精准识别LLM前馈网络中的关键通道,实现激活稀疏性,大幅降低推理成本。
arXiv:2607.27591v1 Announce Type: new Abstract: Feed-forward networks (FFNs) dominate memory traffic and computation in large language model (LLM) inf…