Scaling Self-Play with Self-Guidance
自博弈遇上自指导:新方法以自我引导驱动大规模训练,显著提升智能体策略多样性。
arXiv:2604.20209v2 Announce Type: replace Abstract: LLM self-play algorithms are notable in that, in principle, nothing bounds their learning: a Conje…
自博弈遇上自指导:新方法以自我引导驱动大规模训练,显著提升智能体策略多样性。
arXiv:2604.20209v2 Announce Type: replace Abstract: LLM self-play algorithms are notable in that, in principle, nothing bounds their learning: a Conje…
采样多样性如何左右大模型推理扩展的缩放规律?这项研究给出了新视角。
arXiv:2502.11027v5 Announce Type: replace Abstract: Large language model (LLM) scaling inference is key to unlocking greater performance, and leveragi…
大模型类比能力多样性研究揭示AI推理新维度,或为认知智能评估提供全新视角。
arXiv:2608.03233v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable potential for analogy making, a core cogniti…
探究大模型能否跳出单一推理路径,自主发现多元异构推断,视角新颖。
arXiv:2608.02867v1 Announce Type: new Abstract: Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large l…
前沿大模型在理想化AI竞赛中展露极端策略,而人类决策更趋多元,视角新颖的实验揭示两者差异。
arXiv:2608.01193v1 Announce Type: new Abstract: An AI development race creates a multi-agent safety dilemma. Each company can develop slowly and safel…
硅谷神经多样性招聘为何失灵?LLM正在悄然改写面试规则,让自闭症人才不再被埋没。
Why Silicon Valley's 'Neurodiversity Hiring' Programs Are Failing Autistic Workers (And How LLMs Are Quietly Fixing the Interview) Introduction: The G…
76,636个AI模型实时协作、辩论与批判,揭秘集体智能如何超越个体能力边界
Article URL: https://github.com/ailinone/collective-intelligence Comments URL: https://news.ycombinator.com/item?id=49053465 Points: 7 # Comments: 1
多数投票假设被质疑:这篇研究通过能力控制的审计方法,检验现有多样性指标是否真实衡量了LLM集成中的多样性,结果或有颠覆性。
arXiv:2607.20768v1 Announce Type: cross Abstract: Majority voting over LLMs is widely assumed to benefit from diversity, and diversity measures are us…
注入随机噪声增强LLM推理多样性,ICML 2026 Spotlight论文新思路
arXiv:2605.11936v2 Announce Type: replace Abstract: Recent soft prompt research has tried to improve reasoning by inserting trained vectors into LLM i…
大模型模拟意见时,多样性并不靠堆数量,这项研究揭示了真正影响LLM意见多样性的关键因素
arXiv:2607.20429v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate diverse human opinions in open-ended tasks s…
LLM推理会扼杀创造力?这篇论文揭示AI在游戏中因过度推理导致策略多样性急剧下降的现象。
arXiv:2607.19523v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks, but it…
揭示LLM多样性丧失的深层数据原因,提出Verbalized Sampling方法缓解模式坍塌
arXiv:2510.01171v4 Announce Type: replace Abstract: Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collaps…
元学习攻克多语言LLM对齐难题,让偏好学习适应不同语言文化
arXiv:2607.13315v1 Announce Type: new Abstract: Unequal availability of human preference data across languages poses a significant challenge for align…
Gordon Burtch这本指南直击AI与代码同质化痛点,提供破解单一种植的实用思路。
arXiv:2607.13077v1 Announce Type: cross Abstract: Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assi…
提出Thinking Seeds方法,利用历史多样性与位置感知强化学习,显著提升LLM推理效率与泛化能力。
arXiv:2601.21476v2 Announce Type: replace Abstract: On-policy reinforcement learning (RL) for language model post-training suffers from a fundamental …
从模型多样性切入,LEMUR 2 为AI泛化能力提供新思路,值得算法研究者细读。
arXiv:2607.06839v1 Announce Type: new Abstract: Existing NAS benchmarks (e.g., NAS-Bench, NATS-Bench) cover only narrow, task-specific regions of the …
用大模型生成多样代码提升软件可靠性?这项实证研究给出严谨验证,值得开发者关注。
arXiv:2607.03174v1 Announce Type: cross Abstract: Software diversity has been extensively studied as a means of reducing the risk of common-mode failu…
预测相同、解释却大相径庭,这项研究为可解释机器学习敲响警钟。
arXiv:2512.22240v5 Announce Type: replace-cross Abstract: Machine learning models are primarily judged by predictive performance, especially in applie…
用LLM模拟AI组织中的单极vs多极、隐藏vs可见代理等权衡,一个值得探索的新方向
Currently there doesn't seem to be anyone trying to simulate the tradeoffs between a singleton and a multipolar setup. Or looking at other tradeoffs l…
诊断强化学习训练的Lean定理证明器在推理时的多样性缺陷,为AI数学推理的鲁棒性提供新视角。
arXiv:2601.16172v3 Announce Type: replace Abstract: RL-trained Lean theorem provers mode-collapse at inference time: on miniF2F-test with DeepSeek-Pro…