1
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL
把大模型当作导师,用策略感知的提示自适应破解非可验证强化学习难题,值得关注。
arXiv:2607.04412v1 Announce Type: new Abstract: Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges…