GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding
提出GPTQ-2D算法,实现三次时间复杂度的双向自适应舍入,有效提升大模型量化精度。
arXiv:2607.27042v1 Announce Type: cross Abstract: Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a rea…
提出GPTQ-2D算法,实现三次时间复杂度的双向自适应舍入,有效提升大模型量化精度。
arXiv:2607.27042v1 Announce Type: cross Abstract: Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a rea…
华为Ascend NPU上OpenPangu量化的实证研究,揭秘国产芯片AI模型部署关键优化。
arXiv:2606.21257v1 Announce Type: cross Abstract: OpenPangu models are attractive targets for private and domestic large-language-model deployment, ye…
新方法MosaicQuant通过分离内点与外点实现统一4-bit大模型量化,提升精度与效率。
arXiv:2606.15652v1 Announce Type: new Abstract: 4-bit quantization significantly reduces the memory footprint and accelerates the inference of large l…
提出基于块尺度初始化的NVFP4后训练量化方法,有效提升大语言模型低比特精度。
arXiv:2606.07618v1 Announce Type: new Abstract: NVFP4 is a recently introduced hardware-supported FP4 format that improves the fidelity of 4-bit quant…
仅靠排序即可实现大模型张量化,比传统分解方法更简洁高效,是LLM压缩与加速的新范式
arXiv:2606.08565v1 Announce Type: new Abstract: Tensor networks provide efficient representations for compressing large neural networks. By carefully …
重新审视张量分解在LLM压缩中的核心作用,为后训练压缩提供新思路。
arXiv:2606.03465v1 Announce Type: cross Abstract: Post-training compression is essential for deploying large language models (LLMs) under tight resour…