硅谷今日最热具身模型!不用后训练,看一遍就学会
具身智能迈向GPT时刻
具身智能迈向GPT时刻
AI自主训练AI的真相:执行与迭代被混淆,实证拆解后训练缺失环节
arXiv:2608.19072v1 Announce Type: new Abstract: Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch tr…
用后训练让大模型自动形式化运筹问题,规模化方案值得关注
arXiv:2604.16804v3 Announce Type: replace-cross Abstract: Optimization problems are central to decision-making in manufacturing, logistics, scheduling…
数据选择新思路:用DPO偏好优化为模型量身筛选训练数据,提升后训练效果。
arXiv:2608.16926v1 Announce Type: new Abstract: Data selection in supervised fine-tuning aims to select a small set of effective samples from large-sc…
流形覆盖结合稀疏特征覆盖,为LLM后训练数据筛选提供分层新思路。
arXiv:2608.16927v1 Announce Type: cross Abstract: As supervised fine-tuning data continues to scale, selecting high-value subsets from large candidate…
新研究提出CalibDCD方法,校准LLM后训练特征偏移,显著提升数据污染检测表现,值得关注。
arXiv:2608.10462v1 Announce Type: new Abstract: Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain c…
用MoE代理模型低成本复现LLM强化学习后训练故障,大幅降低排查成本,直击调试痛点。
arXiv:2608.10823v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training of large language models (LLMs) is computationally intensive…
进化策略竟能与GRPO精度持平?揭秘两者在LLM后训练中的不同几何路径。
arXiv:2604.01499v2 Announce Type: replace Abstract: Evolution Strategies (ES) have emerged as a scalable gradient-free alternative to reinforcement le…
超越“可解性”视角,提出“任务可学习性”作为LLM强化学习后训练的静态先验,值得关注。
arXiv:2608.09217v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capa…
没有外部监督的在线自蒸馏方案,为LLM后训练省去人工标注与奖励模型依赖。
arXiv:2608.06296v1 Announce Type: new Abstract: On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language…
针对大模型因果推理短板,探究后训练能否带来质的飞跃,实验结论颇具启发。
arXiv:2602.06337v2 Announce Type: replace-cross Abstract: Causal inference is essential for decision-making but remains challenging for non-experts. W…
协同进化替代反向传播,资源受限环境下多轮工具型智能体的全参数后训练更省内存,性能不打折
arXiv:2608.02391v1 Announce Type: new Abstract: Tool-using large language model (LLM) agents produce long, multi-turn trajectories, making gradient-ba…
Fisher重加权剪枝,让大模型压缩更省能耗,部署更可持续。
arXiv:2608.00481v1 Announce Type: new Abstract: One-shot post-training pruning is the most energy-frugal compression strategy for largelanguage models…
只需8GB显存,就能跑通SFT、DPO、GRPO全流程,轻松理解DeepSeek-R1背后的推理训练奥秘。
Article URL: https://github.com/pochenai/nano-llm-posttraining Comments URL: https://news.ycombinator.com/item?id=49133851 Points: 6 # Comments: 0
结构感知数据组织新方法,直击LLM后训练效率痛点,值得技术党深读
arXiv:2607.27273v1 Announce Type: new Abstract: Post-training of large language models is expensive, and existing efficiency improvements mainly focus…
视觉Transformer部署提速新思路,用脆弱性引导混合精度量化,兼顾压缩率与精度的双赢方案。
arXiv:2607.28589v1 Announce Type: cross Abstract: Post-training quantization (PTQ) has emerged as an effective solution for deploying Vision Transform…
打破LLM后训练算力瓶颈,异构感知奖励引导优化让HPC任务上的强化学习更高效精准。
arXiv:2607.28301v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) can equip large language models (LLMs) with domain knowledge for high-per…
仅用4B参数实现性能反超,中国Neo Lab的持续学习方案让大模型“瘦身”也能拿SOTA。
文|王欣逸 编辑|张雨忻 见到Mind Lab创始人陈锴杰,是在北京的晚上9点半,他已经见了一天的投资人。 陈锴杰是一位连续创业者,从杜克大学休学,做过AI互动故事平台MidReal,也推出了Personal Agent应用Macaron(马卡龙),上线当天就登顶了Product Hunt日榜;20…
Fisher信息加权通道敏感度,为多模态大模型后训练量化提供新高效方案。
arXiv:2607.21076v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) require huge memory and computational costs, which limits the…
欧洲团队开源10B推理语言模型,MIT许可,多语言预训练+强化推理后训练。
arXiv:2607.20448v1 Announce Type: cross Abstract: We introduce Domyn-Small, a 10-billion-parameter open-weight reasoning language model released under…