Scaling Domain Data Repetition in LLM Pretraining
LLM预训练中领域数据重复的缩放规律,揭示数据重复如何影响模型性能与训练效率的边界效应。
arXiv:2608.14071v1 Announce Type: new Abstract: As large language models scale, their training-token budgets must also increase to maintain an appropr…
LLM预训练中领域数据重复的缩放规律,揭示数据重复如何影响模型性能与训练效率的边界效应。
arXiv:2608.14071v1 Announce Type: new Abstract: As large language models scale, their training-token budgets must also increase to maintain an appropr…
揭秘大模型性能增长的底层规律,一文读懂 Scaling Laws 如何指导 LLM 训练与架构设计。
Article URL: https://aidoses.substack.com/p/scaling-laws-the-law-behind-every Comments URL: https://news.ycombinator.com/item?id=49111433 Points: 2 # …
研究发现大模型欺骗行为与预训练语言覆盖度成反比:覆盖越广,模型越少耍花招。
arXiv:2607.24769v1 Announce Type: new Abstract: With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-r…
从零训练30M参数LLM,挑战“缩放地板”假说,实战解析小模型潜力
Article URL: https://github.com/rishipadhye/my-LLM Comments URL: https://news.ycombinator.com/item?id=48997328 Points: 3 # Comments: 0
新研究揭示量化大语言模型的可靠性缩放定律,为部署经济高效且可信赖的AI模型提供理论依据。
arXiv:2607.10855v1 Announce Type: new Abstract: Quantization is a powerful strategy to build capable and resource-efficient large language models (LLM…
从自然语言统计规律出发,揭示神经缩放定律的数学根源,为理解大模型能力增长提供理论基石。
arXiv:2602.07488v3 Announce Type: replace-cross Abstract: Despite the fact that experimental neural scaling laws have substantially guided empirical p…
首个系统研究睡眠基础模型预训练与扩展规律,揭秘数据规模与模型性能的缩放法则,医疗AI从业者必读。
arXiv:2603.00190v2 Announce Type: replace-cross Abstract: Polysomnography (PSG) provides the gold standard for sleep assessment but suffers from subst…
算力持续扩展下,前沿模型能力与小型开发者预算的差距会拉大还是收敛?新研究用两类指标给出答案。
arXiv:2607.00913v1 Announce Type: new Abstract: As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is…
从LLM缩放定律出发,验证传感器数据同样遵循可预测的损失下降规律,为端侧AI和健康监测提供新视角。
Article URL: https://www.empirical.health/blog/llm-scaling-laws-hold-for-sensor-data/ Comments URL: https://news.ycombinator.com/item?id=48741774 Poin…
打破缩放定律瓶颈,用数据高效蒸馏框架低成本训练强推理模型。
arXiv:2508.09883v2 Announce Type: replace Abstract: Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algori…
研究大模型在特定任务蒸馏中缩放定律,揭示蒸馏效率与模型规模、数据量的关系
arXiv:2606.24747v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong performance across a growing range of domains, yet their s…
大模型缩放指数为何“小”?一文拆解算力增长背后的物理极限与效率真相。
arXiv:2606.24504v1 Announce Type: new Abstract: We discuss reasons why the scaling exponents of current Large Language Models (LLMs) applications are …
LLM在科学领域的采用寿命正在缩短,这项研究首次量化了模型过时速度,挑战了传统缩放定律的认知。
arXiv:2604.07530v2 Announce Type: replace-cross Abstract: Scaling laws describe how language model capabilities grow with compute and data, but say no…
揭秘Transformer缩放定律背后的学习动力学与泛化机制,87页长文深度统一理论框架。
arXiv:2512.22088v3 Announce Type: replace-cross Abstract: The scaling law, a cornerstone of Large Language Model (LLM) development, predicts improveme…
揭秘数据混合对模型缩放的影响规律,为AI训练中的最优数据配比提供理论解释。
arXiv:2606.08167v1 Announce Type: new Abstract: Recent research has established empirical scaling laws to predict model performance on multi-domain da…
用潜变量模型解释大模型缩放定律,为架构与基准激增提供理论框架。
arXiv:2512.06553v2 Announce Type: replace-cross Abstract: We propose a statistical framework built on latent variable modeling for scaling laws of lar…
揭示低资源语言模型在多轮、多语言、多阶段训练中的新缩放定律,为高效预训练提供理论指导与实践依据。
arXiv:2410.12325v2 Announce Type: replace Abstract: In this paper, we study a fundamental design problem in pretraining Large Language Models (LLMs) f…
从理论层面揭示LLM训练中高质量数据的最优调度策略,基于质量感知功能缩放定律给出数据使用时机。
arXiv:2605.25698v1 Announce Type: cross Abstract: High-quality data is scarce in large language model (LLM) training, yet how to schedule its use join…
从香农信息论视角重新审视LLM,揭示模型容量与缩放定律的深层联系,ICML 2026前沿研究。
arXiv:2605.23901v1 Announce Type: cross Abstract: Existing scaling laws for Large Language Models (LLMs), predominantly monotonic power laws, fail to …
ICML 2025收录,揭示数据质量如何决定大模型损失与缩放定律的深层关系。
arXiv:2502.12120v3 Announce Type: replace Abstract: Scaling laws guide the development of large language models (LLMs) by offering estimates for the o…