1
GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation
量化分层按输入分级,显著提升LLM生成效率,值得算法工程师细读。
arXiv:2606.23419v1 Announce Type: cross Abstract: Autoregressive decoding with LLMs is primarily bottlenecked by GPU memory bandwidth, especially in e…