RamseyGadgets: A Graph Construction Dataset for LLMs
LLM能像人类一样构造特殊图吗?这个数据集专门来测,为图论+AI交叉研究提供新基准。
arXiv:2608.14999v1 Announce Type: cross Abstract: Constructing special graphs is an important task within graph theory and computer science. Many popu…
LLM能像人类一样构造特殊图吗?这个数据集专门来测,为图论+AI交叉研究提供新基准。
arXiv:2608.14999v1 Announce Type: cross Abstract: Constructing special graphs is an important task within graph theory and computer science. Many popu…
GPT-5.6 霸气出手,中国医生16小时破解22年世界级数学难题,AI推理能力再破纪录!
IT之家 8 月 14 日消息,2025 年,一连串的协和 4+4 博士引发关注和热议,其中就包括金山木事件。时隔一年之后,金山木这个名字再次引发热议。 据《南华早报》今日报道,北京协和医院神经外科博士后、住院医师金山木利用 OpenAI 的 GPT-5.6-Sol 模型,只用了约 16 小时就成功…
顶尖AI模型冲击黎曼猜想,数学难题离破解更近一步,看科技如何改写数学未来。
Article URL: https://www.wsj.com/tech/ai/ai-math-riemann-hypothesis-anthropic-openai-22f98a87 Comments URL: https://news.ycombinator.com/item?id=49306…
新方法让大模型多语言数学推理更强,策略蒸馏带来显著提升,值得一看。
arXiv:2608.05802v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LL…
用数学家的严谨视角拆解AI存在性风险,框架清晰、论证硬核,适合理性派读者。
Article URL: https://alkjash.github.io/ai-risk/ Comments URL: https://news.ycombinator.com/item?id=49161830 Points: 1 # Comments: 0
让大模型自我纠错数学推理,AMTFV用工具流验证把错误答案掰正,值得一试。
arXiv:2607.29549v1 Announce Type: new Abstract: Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably…
OpenAI内部模型推翻80年几何难题,还解决10大开放数学问题,AI推理能力迎来新突破。
OpenAI has publicly reported that an internal general-purpose reasoning model autonomously disproved the Erdős unit distance problem , an open questio…
腾讯混元Hy3模型驱动的科研智能体Hyra突破数学困境,AI助力科学发现再下一城。
IT之家 7 月 31 日消息,腾讯混元今日发文宣布,基于 Hy3 模型的科研智能体 Hyra 找到了一个关键构造, 为加法组合学中一个悬而未决半个多世纪的开放问题给出了完整答案 。 目前,论文预印本、显式构造和形式化证明均已公开: 论文: https://arxiv.org/abs/2607.27…
衡量大模型破解密码的能力,数学推理与网络安全的完美交叉测试。
There’s new benchmark measuring AI’s ability to perform mathematical cryptanalysis. Anthropic’s frontier model actually found new at…
SageMath增强LLM智能体,一场计算与实验数学的自动化探索,看AI如何攻克严谨数学推理。
arXiv:2607.06820v1 Announce Type: new Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, l…
多智能体协作解数学题,用事实图谱记忆减少推理漂移,刷新复杂数学基准表现
arXiv:2607.06447v1 Announce Type: new Abstract: Recent LLM-based mathematical reasoning agents have begun to tackle research-level problems and, in se…
大模型能算对但说不清?新研究解耦数学可解性的潜在方向,直击推理与表达错位
arXiv:2607.05013v1 Announce Type: cross Abstract: Although LLMs have made significant progress in mathematical reasoning, determining whether a mathem…
让大模型在数学推理中越用越聪明:ISM把解题策略沉淀为记忆库,实现持续自我进化,这是ICML数学工作坊的前沿探索。
arXiv:2606.31191v1 Announce Type: new Abstract: We propose Intelligent Schema Memory (ISM), a self-evolving memory-augmented system that improves math…
多模态数学推理新视角:从多样化解题路径中挖掘大模型强化学习潜力,研究思路值得关注。
arXiv:2507.02804v2 Announce Type: replace Abstract: Recent progress in large-scale reinforcement learning (RL) has notably enhanced the reasoning capa…
大模型数学证明迎来新解法:Dafny自动验证加持,让AI每一步推理都可机器校验,从「猜」到「证」的关键跨越。
arXiv:2512.10187v3 Announce Type: replace Abstract: LLMs excel at reasoning, but validating their steps remains challenging. Formal verification offer…
最强LLM面对研究级数学也会自信地犯错,这项研究系统梳理失败模式并给出实证分类,值得关注。
arXiv:2606.24902v1 Announce Type: cross Abstract: The "First Proof" benchmark [1] posed ten research-level mathematics questions to the strongest publ…
本文发现LLM推理中的“悬崖词”——单个token即可导致数学运算失败,揭示模型脆弱性根源。
arXiv:2606.25524v1 Announce Type: new Abstract: Large language models (LLMs) reach high accuracy in mathematical reasoning, but individual traces on t…
专攻乌尔都语数学推理的8B模型,为低资源语言AI推理开辟新路径
arXiv:2606.25568v1 Announce Type: new Abstract: Recent LLMs demonstrate strong mathematical reasoning capabilities, but existing gains rely heavily on…
新论文用未知随机变量问题测试大模型数学推理能力,揭示模型真实推理水平
arXiv:2501.11790v5 Announce Type: replace-cross Abstract: Recent studies have raised significant concerns regarding the reliability of current mathema…