1
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents
用双模拟和动作反事实估计突破组相对策略优化的瓶颈,为LLM智能体长程决策提供新思路。
arXiv:2606.25556v1 Announce Type: cross Abstract: Stepwise group-based RL is an attractive way to train long-horizon LLM agents without a learned crit…