1
Show HN: Run GLM-4.5-Air(110B)on a 16GBRAM consumer machine
16GB内存就能运行110B大模型,量化技术让消费级硬件也能玩转前沿AI。
Article URL: https://github.com/FedericoTs/quantprobe Comments URL: https://news.ycombinator.com/item?id=49028865 Points: 1 # Comments: 0
16GB内存就能运行110B大模型,量化技术让消费级硬件也能玩转前沿AI。
Article URL: https://github.com/FedericoTs/quantprobe Comments URL: https://news.ycombinator.com/item?id=49028865 Points: 1 # Comments: 0
提出专家引导的后合并量化方法,利用合并权重锚定,在低资源部署中平衡模型压缩与性能。
arXiv:2605.16882v1 Announce Type: new Abstract: Low-resource deployment constraints have made model quantization essential for deploying neural networ…
1-bit量化大模型新思路,输出对齐策略再审视,助力低资源设备高效推理
arXiv:2512.21651v3 Announce Type: replace Abstract: Large Language Models (LLMs) deliver strong performance across a wide range of NLP tasks, but thei…