1
Show HN: Run GLM-4.5-Air(110B)on a 16GBRAM consumer machine
16GB内存就能运行110B大模型,量化技术让消费级硬件也能玩转前沿AI。
Article URL: https://github.com/FedericoTs/quantprobe Comments URL: https://news.ycombinator.com/item?id=49028865 Points: 1 # Comments: 0
16GB内存就能运行110B大模型,量化技术让消费级硬件也能玩转前沿AI。
Article URL: https://github.com/FedericoTs/quantprobe Comments URL: https://news.ycombinator.com/item?id=49028865 Points: 1 # Comments: 0
面向AMD NPU的融合混合精度内核库,显著提升量化大语言模型推理效率。
arXiv:2606.11357v1 Announce Type: cross Abstract: With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improv…