Scaling Self-Play with Self-Guidance
自博弈遇上自指导:新方法以自我引导驱动大规模训练,显著提升智能体策略多样性。
arXiv:2604.20209v2 Announce Type: replace Abstract: LLM self-play algorithms are notable in that, in principle, nothing bounds their learning: a Conje…
自博弈遇上自指导:新方法以自我引导驱动大规模训练,显著提升智能体策略多样性。
arXiv:2604.20209v2 Announce Type: replace Abstract: LLM self-play algorithms are notable in that, in principle, nothing bounds their learning: a Conje…
自博弈评判者打分的是“像不像”而非“对不对”,参考无关LLM裁判存在结构性缺陷。
arXiv:2607.05904v1 Announce Type: new Abstract: Training a language model against its own reference-free judgments (the premise of self-rewarding, sel…
自博弈训练让AI修复代码能力再进一步,锚定策略稳定训练过程,ICML论文值得读。
arXiv:2607.03523v1 Announce Type: cross Abstract: Code repair is an important capability for language models (LMs): given a buggy program and unit tes…
提出一种开放式自我改进推理器OpenSIR,通过自博弈突破标注数据限制,有望超越人类水平推理。
arXiv:2511.00602v3 Announce Type: replace Abstract: Recent advances in large language model (LLM) reasoning through reinforcement learning rely on ann…
用树状自博弈让代码LLM学会避免常见漏洞,安全对齐新思路。
arXiv:2606.03489v1 Announce Type: cross Abstract: While Large Language Models (LLMs) excel in code generation, they remain prone to replicating subtle…