Show HN: Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO)
只需8GB显存,就能跑通SFT、DPO、GRPO全流程,轻松理解DeepSeek-R1背后的推理训练奥秘。
Article URL: https://github.com/pochenai/nano-llm-posttraining Comments URL: https://news.ycombinator.com/item?id=49133851 Points: 6 # Comments: 0