SchurQuant: Groupwise Discrete Optimization for Layer-Wise LLM Quantization
大模型2-3比特量化精度崩坏?这项研究用分组离散优化突破极限,无需反向传播即可部署。
arXiv:2608.15567v1 Announce Type: new Abstract: Weight-only post-training quantization (PTQ) enables the deployment of large language models under tig…