Training nGPT
把 Transformer 参数与激活约束到超球面,nGPT 带来表示学习新范式,一文看懂归一化训练的关键设计。
arXiv:2608.01284v1 Announce Type: cross Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining mo…
把 Transformer 参数与激活约束到超球面,nGPT 带来表示学习新范式,一文看懂归一化训练的关键设计。
arXiv:2608.01284v1 Announce Type: cross Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining mo…
揭秘大脑连续工作记忆的数学机制,看归一化如何塑造低秩慢流形。
arXiv:2608.01947v1 Announce Type: cross Abstract: The ability to robustly maintain and update continuous variables is a hallmark of working memory. Wh…
提出归一化奖励方法,提升偏好优化训练稳定性与效果
arXiv:2607.16240v1 Announce Type: cross Abstract: Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs…
口音归一化新研究,离散扩散模型实现可控语音转换,论文思路前沿,值得语音方向开发者深挖。
arXiv:2603.14275v2 Announce Type: replace-cross Abstract: Existing accent normalization methods do not typically offer control over accent strength, y…
用全局归一化稳定多模态大模型基于策略的蒸馏,提升推理性能与训练效率的创新方法。
arXiv:2606.09091v1 Announce Type: new Abstract: On-policy distillation (OPD) has recently emerged as an important post-training paradigm. By using a s…
突破PreNorm与PostNorm的困境:SpanNorm在提升深度Transformer训练稳定性同时保持高性能
arXiv:2601.22580v2 Announce Type: replace-cross Abstract: The success of Large Language Models (LLMs) hinges on the stable training of deep Transforme…
通过pass-rate加权自蒸馏,恢复LLM推理的“甜蜜点”,破解GRPO归一化带来的学习偏差。
arXiv:2605.27765v1 Announce Type: cross Abstract: Self-Distillation Policy Optimization (SDPO) provides dense token-level credit assignment for reinfo…
提出归一化等变性的结构先验,可应用于任意骨干网络的图像去噪,有效提升分布偏移健壮性。
arXiv:2605.08193v2 Announce Type: replace-cross Abstract: Normalization Equivariance (NE) is a structural prior that improves robustness to distributi…