Scaling Domain Data Repetition in LLM Pretraining
LLM预训练中领域数据重复的缩放规律,揭示数据重复如何影响模型性能与训练效率的边界效应。
arXiv:2608.14071v1 Announce Type: new Abstract: As large language models scale, their training-token budgets must also increase to maintain an appropr…
LLM预训练中领域数据重复的缩放规律,揭示数据重复如何影响模型性能与训练效率的边界效应。
arXiv:2608.14071v1 Announce Type: new Abstract: As large language models scale, their training-token budgets must also increase to maintain an appropr…
采样多样性如何左右大模型推理扩展的缩放规律?这项研究给出了新视角。
arXiv:2502.11027v5 Announce Type: replace Abstract: Large language model (LLM) scaling inference is key to unlocking greater performance, and leveragi…
大模型剪枝新招:高层注意力剪枝加缩放,推理提速不损精度,论文原文速看。
arXiv:2507.01900v3 Announce Type: replace-cross Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), signifi…
别只调温度了!双层优化为LLM校准带来更精准的建模方案。
arXiv:2608.07419v1 Announce Type: new Abstract: Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. Tra…
视频大模型缩放规律新发现:稳定曲线下藏着不稳定的样本差异,读懂这项研究才能避开评估陷阱。
arXiv:2608.07014v1 Announce Type: new Abstract: Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. …
揭秘大模型性能增长的底层规律,一文读懂 Scaling Laws 如何指导 LLM 训练与架构设计。
Article URL: https://aidoses.substack.com/p/scaling-laws-the-law-behind-every Comments URL: https://news.ycombinator.com/item?id=49111433 Points: 2 # …
研究发现大模型欺骗行为与预训练语言覆盖度成反比:覆盖越广,模型越少耍花招。
arXiv:2607.24769v1 Announce Type: new Abstract: With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-r…
IT之家 7 月 28 日消息,科技媒体 Android Authority 昨日(7 月 27 日)发布博文,报道称在 Galaxy Z Fold8 等折叠手机上, 三星正酝酿精细化应用缩放功能,支持用户为某个应用设置 5 档缩放级别。 该功能目前仍处于开发阶段,路径位于“设置 → 高级功能 → …
新方法通过奖励感知群体缩放优化进化策略,显著提升LLM微调效率,已被ICML 2026 workshop收录。
arXiv:2607.19408v1 Announce Type: new Abstract: Using Evolutionary Strategies (ES) for fine-tuning large language models is attractive because it is m…
从零训练30M参数LLM,挑战“缩放地板”假说,实战解析小模型潜力
Article URL: https://github.com/rishipadhye/my-LLM Comments URL: https://news.ycombinator.com/item?id=48997328 Points: 3 # Comments: 0
新研究揭示量化大语言模型的可靠性缩放定律,为部署经济高效且可信赖的AI模型提供理论依据。
arXiv:2607.10855v1 Announce Type: new Abstract: Quantization is a powerful strategy to build capable and resource-efficient large language models (LLM…
一篇探索生成式AI放大效应预测的前沿论文,从理论框架到实验验证均有突破性贡献。
arXiv:2509.08048v4 Announce Type: replace-cross Abstract: Generative networks are perfect tools to enhance the speed and precision of LHC simulations.…
用模糊推理动态调整私有区块链验证节点数量,解决资源浪费和性能瓶颈。
arXiv:2607.07901v1 Announce Type: cross Abstract: Private blockchain networks run with fixed node configurations that cannot adapt to changing workloa…
从自然语言统计规律出发,揭示神经缩放定律的数学根源,为理解大模型能力增长提供理论基石。
arXiv:2602.07488v3 Announce Type: replace-cross Abstract: Despite the fact that experimental neural scaling laws have substantially guided empirical p…
首个系统研究睡眠基础模型预训练与扩展规律,揭秘数据规模与模型性能的缩放法则,医疗AI从业者必读。
arXiv:2603.00190v2 Announce Type: replace-cross Abstract: Polysomnography (PSG) provides the gold standard for sleep assessment but suffers from subst…
字节跳动开源EdgeBench,揭示AI在真实世界环境中的学习缩放规律,值得研究者关注。
Article URL: https://github.com/ByteDance-Seed/EdgeBench Comments URL: https://news.ycombinator.com/item?id=48801308 Points: 1 # Comments: 0
探讨LLM缩放能否提升社会模拟的保真度,揭示当前范式的局限与挑战
arXiv:2607.02464v1 Announce Type: new Abstract: Large Language Model (LLM) social simulations are a promising research method, but they are not yet fa…
算力持续扩展下,前沿模型能力与小型开发者预算的差距会拉大还是收敛?新研究用两类指标给出答案。
arXiv:2607.00913v1 Announce Type: new Abstract: As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is…
从LLM缩放定律出发,验证传感器数据同样遵循可预测的损失下降规律,为端侧AI和健康监测提供新视角。
Article URL: https://www.empirical.health/blog/llm-scaling-laws-hold-for-sensor-data/ Comments URL: https://news.ycombinator.com/item?id=48741774 Poin…
捏合缩放代替逐层点击,树形导航的全新交互原型值得一试
I got this idea recently when I had to browse files more than usual and felt the pain: why do I need to click through every level when I want to drill…