Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay
一步错步步错?实测三类信用信号在LLM智能体训练中均无法锁定关键步骤,揭示因果归因盲区。
arXiv:2608.19760v1 Announce Type: new Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld)…