Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling
多角色评分标准生成,让LLM裁判与奖励模型更稳健、更对齐人类偏好。
arXiv:2607.01830v1 Announce Type: new Abstract: Reliable reward and preference signals are critical for evaluating and optimizing large language model…