1
AirLLM Runs a 70B Model on a 4GB GPU. It's True, and That's Not the Interesting Part
70B大模型跑进4GB显存,不靠量化保住全精度,原理比结果更惊艳
AirLLM's README opens with a line that sounds like it can't be true: AirLLM dramatically reduces inference memory usage, letting 70B large language mo…