Projector Is All You Train
投影器成训练核心?这篇论文挑战传统微调范式,或为高效训练提供新思路。
arXiv:2608.19726v1 Announce Type: cross Abstract: The typical training process of a multimodal large language model (MLLM) involves adapting both the …
投影器成训练核心?这篇论文挑战传统微调范式,或为高效训练提供新思路。
arXiv:2608.19726v1 Announce Type: cross Abstract: The typical training process of a multimodal large language model (MLLM) involves adapting both the …
用烘焙类比LLM训练,把复杂技术讲得色香味俱全,看完想动手“烤”一个模型。
Article URL: https://newsletter.kentbeck.com/p/baking-a-model Comments URL: https://news.ycombinator.com/item?id=49305969 Points: 2 # Comments: 0
用强化学习给AI数据中心“降耗”,实测从单卡到集群的LLM训练功率控制,兼顾性能与能效。
arXiv:2608.11226v1 Announce Type: new Abstract: Reinforcement-learning post-training dominates modern language-model development, yet its power behavi…
大模型预训练遇卡顿不知根源?SCOUT用共识异常检测精准锁定故障节点,摆脱同步困局。
arXiv:2608.11034v1 Announce Type: cross Abstract: In LLM pre-training, synchronization propagates rank-local stalls, slowdowns, and numerical errors i…
全新模块化更新解耦机制,让大模型训练内存更省、并发更强,值得关注的前沿方案。
arXiv:2608.07974v1 Announce Type: new Abstract: Large language model (LLM) fine-tuning at the edge adapts the model to scenario-specific data while pr…
人文社科数据稀缺?这项研究提出HSS-Synth合成方案,专攻开放场景下的大模型数据难题。
arXiv:2607.27379v1 Announce Type: new Abstract: High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly. Da…
用稠密层次奖励与课程学习提升代码大模型训练效果,为AI编程能力突破提供新路径。
arXiv:2607.26457v1 Announce Type: new Abstract: Reinforcement learning is a natural post-training paradigm for code-oriented large language models bec…
海上核能数据中心破解AI算力能源困局,探讨核动力浮式数据中心新方案
Architecting Maritime Nuclear Micro-Reactors for Data Center Infrastructure The exponential growth of large language model (LLM) training and high-per…
类似加密挖矿的志愿者池化训练,能否让开源社区共同训练前沿大模型?
Inspired by Crypto Pool Mining: would it be possible to train LLMs as a pool of volunteers? I feel like that would be a good way to have frontier mode…
从零训练30M参数LLM,挑战“缩放地板”假说,实战解析小模型潜力
Article URL: https://github.com/rishipadhye/my-LLM Comments URL: https://news.ycombinator.com/item?id=48997328 Points: 3 # Comments: 0
针对大模型训练中的故障恢复痛点,提出软硬件协同设计的开源蓝图,值得关注
Article URL: https://github.com/PJHkorea/pim-hbm-bypass Comments URL: https://news.ycombinator.com/item?id=48940625 Points: 2 # Comments: 1
无需隐变量,三值量化LLM训练新方法,开源项目BitBop让你以极低成本体验高效训练。
Article URL: https://github.com/ValerioDolci/bitbop Comments URL: https://news.ycombinator.com/item?id=48892013 Points: 1 # Comments: 1
NVIDIA官方详解如何利用主机卸载技术,在JAX中缓解HBM瓶颈,支持更大规模LLM训练。
Article URL: https://developer.nvidia.com/blog/reducing-high-bandwidth-memory-bottlenecks-in-jax-based-llm-training-with-host-offloading/ Comments URL…
教你用记忆化指导数据重用的策略,在不大幅延长训练时间的前提下提升LLM训练效率,顶级会议论文的实用思路。
arXiv:2607.04969v1 Announce Type: new Abstract: The training paradigm of large language models has shifted from traditional one-pass training to multi…
零开销检查点热替换技术,让大模型训练可无缝应对软硬件故障,大幅提升训练韧性
arXiv:2607.01646v1 Announce Type: new Abstract: State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing un…
融合工业控制理论,用卡尔曼滤波+PID+史密斯预测器精准防止LLM训练时GPU OOM,开源实现值得一试。
Article URL: https://github.com/sajjaddoda72-design/UATC Comments URL: https://news.ycombinator.com/item?id=48734992 Points: 2 # Comments: 0
揭秘LLM训练中学习率缩放的非线性规律,为优化大模型训练过程提供全新理论视角。
arXiv:2606.29158v1 Announce Type: new Abstract: Learning-rate transfer can reduce the cost of training large language models: instead of sweeping lear…
基于几何原理的随机优化方法,为大规模语言模型训练效率提供新思路
arXiv:2510.01878v2 Announce Type: replace Abstract: Low-rank gradient optimization for large language models is currently divided into two categories:…
提出机制驱动监控器,抢先检测LLM训练不稳定,提升大模型训练可靠性
arXiv:2606.28116v1 Announce Type: new Abstract: Frontier large language model training consumes massive accelerator fleets and long wall-clock computa…
融合反向传播与优化器阶段的梯度处理,大幅降低LLM训练的内存峰值,突破显存瓶颈的新方案。
arXiv:2606.22932v1 Announce Type: new Abstract: Reverse-mode differentiation computes every weight gradient, writes it to memory, and only then lets t…