1
QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
量化压缩不再掉精度,按token动态分配层计算,推理效率与质量兼得。
arXiv:2602.10431v4 Announce Type: replace Abstract: Large language models (LLMs) demand substantial computational and memory resources, posing challen…