GPT-Red: Automated Red Teaming via Self-Play at Scale
提出GPT-Red方法,通过大规模自对弈自动生成对抗性测试,提升大模型安全性。
arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prom…
提出GPT-Red方法,通过大规模自对弈自动生成对抗性测试,提升大模型安全性。
arXiv:2607.26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prom…
ICML 2026收录,S-SPPO用语义校准提升自对弈偏好优化,为AI对齐训练提供新思路。
arXiv:2606.01561v1 Announce Type: cross Abstract: Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preferen…
自对弈算法在定理证明领域有了严谨的理论根基,为AI数学推理提供新视角。
arXiv:2606.01861v1 Announce Type: new Abstract: Self-play, a type of training algorithm that enables a model to self-improve, has recently shown promi…
大模型自对弈新范式:群体进化生成可解任务,驱动推理能力持续攀升
Article URL: https://vmax.ai/team/populora-co-evolving-llm-populations-for-reasoning-self-play Comments URL: https://news.ycombinator.com/item?id=4821…