1
Kog is going deeper to squeeze more inference out of GPUs
Kog专攻现有数据中心GPU的极致推理提速,演示单请求极速解码,还拿下200条销售线索,值得关注。
The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog.
Kog专攻现有数据中心GPU的极致推理提速,演示单请求极速解码,还拿下200条销售线索,值得关注。
The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog.
70B大模型跑进4GB显存,不靠量化保住全精度,原理比结果更惊艳
AirLLM's README opens with a line that sounds like it can't be true: AirLLM dramatically reduces inference memory usage, letting 70B large language mo…