Matching Accuracy, Different Geometry: Evolution Strategies vs GRPO in LLM Post-Training
进化策略竟能与GRPO精度持平?揭秘两者在LLM后训练中的不同几何路径。
arXiv:2604.01499v2 Announce Type: replace Abstract: Evolution Strategies (ES) have emerged as a scalable gradient-free alternative to reinforcement le…