TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents
面向长时程智能体,TRCA用过渡级评分信用分配破解稀疏奖励难题,值得关注。
arXiv:2608.16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, …