LLM City – 3D render of all Kimi K3's weights as 2.5mm tiles
把Kimi K3的权重渲染成3D城市:24.9万个张量、896个专家一目了然,AI内部结构从未如此直观。
Article URL: https://magik.net/llmcity/ Comments URL: https://news.ycombinator.com/item?id=49333151 Points: 2 # Comments: 1
把Kimi K3的权重渲染成3D城市:24.9万个张量、896个专家一目了然,AI内部结构从未如此直观。
Article URL: https://magik.net/llmcity/ Comments URL: https://news.ycombinator.com/item?id=49333151 Points: 2 # Comments: 1
MoE训练新方案:Sinkhorn梯度下降替代AdamW,减少优化器状态内存,让大模型训练更省显存。
arXiv:2608.04407v1 Announce Type: cross Abstract: Memory-efficient matrix optimizers such as Sinkhorn gradient descent remove most AdamW optimizer sta…
月之暗面发布2.8万亿参数MoE开源模型,百万token上下文与视觉能力直指前沿闭源模型。
Heads up: This article was written by AgentOne Research . Moonshot AI has officially launched Kimi K3 , a 2.8-trillion-parameter mixture-of-experts (M…
新论文提出UMoE方法,在领域特定训练中激活每个专家,突破传统MoE路径选择限制,提升模型适应性与效率。
arXiv:2607.11444v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale capacity without proportional compute cost and have become a key…
研究用冻结LLM知识做文生图,Mixture-of-Transformers架构实现知识迁移,只靠标准图文对训练。
arXiv:2606.29013v1 Announce Type: new Abstract: Leveraging capabilities of large language models (LLMs) in text-to-image (T2I) synthesis is an importa…
探究 67 个前沿模型组合的边界:路由、投票与 MoA 的共失败上限揭示何时组合无益。
arXiv:2606.27288v1 Announce Type: new Abstract: Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to b…
大模型调优新思路:用混合深度集成技术提升语言模型性能,值得关注!
arXiv:2410.13077v2 Announce Type: replace-cross Abstract: Transformer-based Large Language Models (LLMs) traditionally rely on final-layer loss for fi…
提出ReM-MoA机制,用推理记忆解决混合代理规模化难题,提升多智能体协作效率。
arXiv:2606.24437v1 Announce Type: new Abstract: Mixture-of-Agents (MoA) architectures improve inference-time scaling by organizing multiple LLM agents…
提出SkillMoV混合视图路由架构,用原型条件门控统一多视角能力估计,创新融合视图选择与专家路由
arXiv:2606.17615v1 Announce Type: cross Abstract: Estimating human proficiency from video is a key challenge for automated skill assessment, with appl…
MoE新范式:跨OLMoE、Qwen3等主流架构验证的Tied Expert Layers方法,预训练实验揭示连接专家层如何提升大模型效率。
arXiv:2606.16825v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures efficiently scale Large Language Models (LLMs) by activating …
持续学习新突破:提出一致性保持混合专家模型,有效解决灾难性遗忘问题。
arXiv:2605.20247v1 Announce Type: new Abstract: Catastrophic forgetting remains a major obstacle to continual learning in large language models (LLMs)…
ICML 2026 顶会论文:深入 Mixture-of-Experts 语言模型的专家级别内部机制,揭示专家如何协同与对抗。
arXiv:2604.02178v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) architectures have become the dominant choice for scaling Large Lan…
让AI智能体拥有记忆与递归能力,MMoA框架通过记忆混合智能体显著提升大模型输出质量。
arXiv:2605.19194v1 Announce Type: new Abstract: The Mixture-of-Agents (MoA) framework has shown promise in improving large language model (LLM) perfor…
开源权重模型Kimi K2.6在编程基准测试中力压Claude、GPT-5.5等前沿模型,MoE架构与技术细节值得关注。
The benchmark leaderboard for large language models just shifted again. Moonshot AI's Kimi K2.6, an open-weights model, outperformed Claude, GPT-5.5, …
新框架直接优化群体级专家平衡,解决MoE训练中负载均衡偏差问题。
arXiv:2605.15403v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models rely on balanced expert utilization to fully realize their scalability…