1
ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents
工具型智能体训练新思路:反射增强策略优化,值得关注的前沿方法。
arXiv:2606.31392v1 Announce Type: new Abstract: Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling externa…