1
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
生成式大模型蒸馏奖励模型新方法,解决偏好标注昂贵难题。
arXiv:2601.14032v2 Announce Type: replace Abstract: Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human prefer…