AI 竞赛升温:曝 OpenAI 已预训练 Bel 模型:超 10T 参数,冲击通用人工智能
OpenAI新模型Bel曝光,超10万亿参数直指AGI,AI竞赛再升级。
IT之家 8 月 26 日消息,消息源 @synthwavedd 今天(8 月 26 日)在 X 平台发布推文,爆料称 OpenAI 已完成预训练下一代模型, 内部代号为 Bel,是 Doug 模型的直接继任者,总参数量超过 10T。 消息称 OpenAI 希望将 Bel 模型打造成为 GPT-6 …
OpenAI新模型Bel曝光,超10万亿参数直指AGI,AI竞赛再升级。
IT之家 8 月 26 日消息,消息源 @synthwavedd 今天(8 月 26 日)在 X 平台发布推文,爆料称 OpenAI 已完成预训练下一代模型, 内部代号为 Bel,是 Doug 模型的直接继任者,总参数量超过 10T。 消息称 OpenAI 希望将 Bel 模型打造成为 GPT-6 …
从星系到语言模型,AstroPT揭示了预训练模型如何习得结构知识,为LLM研究提供跨学科洞见。
arXiv:2608.22614v1 Announce Type: new Abstract: Interpretability research increasingly asks when concepts emerge during training and whether linear pr…
巨型概念人形机器人ZERO惊艳亮相,基于NVIDIA Isaac平台预训练,开启未来机器人新想象。
中坚科技(002779.SZ)旗下桦之坚携概念人形机器人ZERO亮相2026世界机器人大会。
用继续预训练给本地小模型注入新领域知识,实战示例清晰,适合想深入LLM调优的读者。
Article URL: https://www.teachmecoolstuff.com/viewarticle/teaching-a-local-llm-a-new-domain Comments URL: https://news.ycombinator.com/item?id=4938012…
揭秘大模型在持续预训练中如何习得、保留与遗忘概念,理解AI认知的关键研究。
arXiv:2601.03570v2 Announce Type: replace Abstract: Human beings primarily understand the world through concepts (e.g., dog), abstract mental represen…
面向代码智能的预训练新范式,用不变学习增强代码表示的鲁棒性与泛化能力,值得关注。
arXiv:2608.15412v1 Announce Type: cross Abstract: Encoder-based code representation models remain widely deployed for discriminative tasks such as clo…
LLM预训练中领域数据重复的缩放规律,揭示数据重复如何影响模型性能与训练效率的边界效应。
arXiv:2608.14071v1 Announce Type: new Abstract: As large language models scale, their training-token budgets must also increase to maintain an appropr…
用测试增强学习替代海量语料,为语言模型持续预训练省下高额计算成本,稳准快。
arXiv:2608.11788v1 Announce Type: cross Abstract: Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large lang…
预训练与压缩小模型谁更可信?IJCNN权威基准测试给出对比答案。
arXiv:2608.11981v1 Announce Type: new Abstract: Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Languag…
用谱视角剖析大模型领域适应,Diffract方法为持续预训练提供新洞察。
arXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language model…
医疗影像中选预训练模型总靠试错?本文系统检验了迁移性估计指标在医学场景下的鲁棒性,帮你少走弯路。
arXiv:2608.09999v1 Announce Type: cross Abstract: In transfer learning, the choice of source model largely influences the performance on a target data…
直击扩散语言模型预训练与生成阶段的不匹配痛点,提出改进方法,值得算法研究者细读。
arXiv:2608.09424v1 Announce Type: new Abstract: Autoregressive language models align training and use: generation conditions on a clean prompt, and tr…
多语言预训练并非总是有益,这项研究揭示了LLM预训练不稳定的关键发现。
arXiv:2608.08800v1 Announce Type: new Abstract: Pretraining LLMs on artificial languages ("pre-pretraining") is a technique that could reportedly incr…
预训练模型暴露如何放大微调大模型的越狱风险,安全研究必读。
arXiv:2512.14751v3 Announce Type: replace-cross Abstract: Finetuning pretrained large language models (LLMs) has become the standard paradigm for deve…
把多模态健康数据与自编码器结合,交叉掩码预训练让数字医疗测量更精准,值得算法研究者一读
arXiv:2506.02260v4 Announce Type: replace-cross Abstract: Wearable devices enable continuous multi-modal physiological and behavioral monitoring, yet …
低资源场景下语音大模型的数据需求,以及高资源语言预训练的影响,实验数据扎实,值得一读。
arXiv:2508.05149v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-…
系统评估20余种初始化策略,为LLM词汇扩展提供实用指南。
arXiv:2608.03494v1 Announce Type: cross Abstract: Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new lan…
用调用图预训练提升二进制分析任务,揭秘上下文如何带来增益,逆向工程研究者值得一看。
arXiv:2608.02084v1 Announce Type: cross Abstract: Binary function embedding models are trained to encode the semantics of binary code in such a way th…
视觉编码器新思路:利用大语言模型进行分层预训练,提升视觉特征学习效率,CVPR 收录论文值得关注。
arXiv:2604.00086v2 Announce Type: replace-cross Abstract: The field of computer vision has experienced significant advancements through scalable visio…