Tokenizer-Agnostic Engram Module
突破传统分词限制,提出与分词器无关的记忆模块,为语言模型的长效记忆与跨词元建模提供新思路。
arXiv:2607.29065v1 Announce Type: new Abstract: Deepseek's Engram, a conditional memory module, was introduced to trade-off storage versus reasoning i…
突破传统分词限制,提出与分词器无关的记忆模块,为语言模型的长效记忆与跨词元建模提供新思路。
arXiv:2607.29065v1 Announce Type: new Abstract: Deepseek's Engram, a conditional memory module, was introduced to trade-off storage versus reasoning i…
燧原科技携手中兴通讯,自研AI芯片与架构创新打造云燧超节点,大幅降低互联成本,引领智算新浪潮。
IT之家 7 月 19 日消息,7 月 18 日下午,在 2026 世界人工智能大会期间,燧原科技正式发布云燧 ESL64-O 超节点、CoPoS+AI 芯片、天基算力应用场景三项成果。 首先,燧原科技联合中兴通讯正式发布 云燧 ESL64-O 超节点 ,号称“以四大能力重新定义从 Bit 到 To…
上海AI Lab以397B参数实现万亿级模型性能,首创记忆与思考解耦架构,推理效率暴增4倍,改写大模型进化逻辑。
采用非Transformer架构“Mobius”,不做泛化通用问答
用强化学习实现分词端到端训练,挑战LLM中最后的硬编码压缩步骤
arXiv:2602.13940v2 Announce Type: replace Abstract: Tokenization is a hardcoded compression step which remains in the training pipeline of Large Langu…
dMoE提出可学习块专家机制,为大型语言模型混合专家设计提供新思路,架构简洁高效。
arXiv:2605.30876v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregres…
DeepSeek V4 Pro激进架构引发价格战,输入输出成本降至对手1/7和1/17,直接冲击硅谷大模型护城河。
DeepSeek’s announcement over the weekend that it has made its 75% price cut permanent on its flagship V4 Pro model is a disruptive assault on the capi…
突破性语言模型架构,用句子曲线提升表达效率与连贯性
arXiv:2602.01807v3 Announce Type: replace-cross Abstract: Language models (LMs) are a central component of modern AI systems, and diffusion language m…
聚焦实时流VLA架构创新,重新思考并加速Flow VLA推理效率,适合AI研究者。
arXiv:2603.19199v3 Announce Type: replace-cross Abstract: Real-time execution is crucial for deploying Vision-Language-Action (VLA) models in the phys…
颠覆传统AI代理架构,抛弃LLM+向量存储,提出全新方案,值得技术研究者一读。
Article URL: https://sbarron.com/writing/substrate-is-the-body Comments URL: https://news.ycombinator.com/item?id=48198625 Points: 2 # Comments: 1