1
REAR: Test-time Preference Realignment through Reward Decomposition
提出REAR方法,在推理阶段通过奖励分解实现偏好重新对齐,为大模型对齐提供高效新解法。
arXiv:2606.30339v1 Announce Type: cross Abstract: Aligning large language models (LLMs) with diverse user preferences is a critical yet challenging ta…