1
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
面向多轮智能体优化,提出前缀感知内部奖励模型,解决长对话稀疏奖励难题。
arXiv:2605.17877v2 Announce Type: replace Abstract: A significant hurdle for current LLMs is the execution of complex, multi-stage tasks. Group Relati…