1
Scaling Self-Play with Self-Guidance
自博弈遇上自指导:新方法以自我引导驱动大规模训练,显著提升智能体策略多样性。
arXiv:2604.20209v2 Announce Type: replace Abstract: LLM self-play algorithms are notable in that, in principle, nothing bounds their learning: a Conje…
自博弈遇上自指导:新方法以自我引导驱动大规模训练,显著提升智能体策略多样性。
arXiv:2604.20209v2 Announce Type: replace Abstract: LLM self-play algorithms are notable in that, in principle, nothing bounds their learning: a Conje…
提出基于闭环再入神经系统的安全AGI架构,直指前馈网络缺乏自指能力的根本缺陷,为下一代智能体安全性提供全新蓝图。
arXiv:2606.26406v1 Announce Type: cross Abstract: We propose a complete architectural blueprint for safe artificial general intelligence based on a cl…