Scaling Laws: The Law Behind Every LLM
揭秘大模型性能增长的底层规律,一文读懂 Scaling Laws 如何指导 LLM 训练与架构设计。
Article URL: https://aidoses.substack.com/p/scaling-laws-the-law-behind-every Comments URL: https://news.ycombinator.com/item?id=49111433 Points: 2 # …
揭秘大模型性能增长的底层规律,一文读懂 Scaling Laws 如何指导 LLM 训练与架构设计。
Article URL: https://aidoses.substack.com/p/scaling-laws-the-law-behind-every Comments URL: https://news.ycombinator.com/item?id=49111433 Points: 2 # …
从零开始的跨模态预训练规模扩展,揭示原生多模态模型的新范式与核心挑战。
arXiv:2607.22043v1 Announce Type: new Abstract: Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on tex…
大模型缩放指数为何“小”?一文拆解算力增长背后的物理极限与效率真相。
arXiv:2606.24504v1 Announce Type: new Abstract: We discuss reasons why the scaling exponents of current Large Language Models (LLMs) applications are …
一种全新训练范式,通过反馈条件优化让LLM在所有训练阶段都获得更好的缩放能力,可能颠覆大模型训练方式。
arXiv:2605.20285v1 Announce Type: new Abstract: We tackle the question of how to scale more efficiently across the many, ever-growing stages of curren…
重磅研究:代码领域的缩放定律显示需要比自然语言多几个数量级的数据才能达到相同性能提升,引发对大模型训练数据效率的重新思考。
arXiv:2510.08702v2 Announce Type: replace Abstract: Code Large Language Models (LLMs) are revolutionizing software engineering. However, scaling laws …
首次揭示LLM代理系统中技能缩放定律,基于15个模型、超百万数据点的实证研究
arXiv:2605.16508v1 Announce Type: new Abstract: As agent systems scale, skills accumulate into large reusable libraries, yet their scaling laws remain…
受限数据下混合预训练的缩放定律,揭示稀缺目标数据与通用数据的最佳配比策略。
arXiv:2605.12715v2 Announce Type: replace Abstract: As language models scale, the amount of data they require grows -- yet many target data sources, s…