AI at Home Part 2: Multi-GPU Drifting
手把手拆解Transformer注意力机制与多GPU并行,家庭AI集群实战避坑指南。
Article URL: https://jdagostino.github.io/ai-pt2-multi-gpu-drifting/index.html Comments URL: https://news.ycombinator.com/item?id=49377155 Points: 2 #…
手把手拆解Transformer注意力机制与多GPU并行,家庭AI集群实战避坑指南。
Article URL: https://jdagostino.github.io/ai-pt2-multi-gpu-drifting/index.html Comments URL: https://news.ycombinator.com/item?id=49377155 Points: 2 #…
多视角眼底图像结合新型视觉Transformer,为卒中快速筛查提供高精度AI方案
arXiv:2608.14722v1 Announce Type: new Abstract: Stroke remains a leading cause of mortality and morbidity worldwide, emphasizing the importance of its…
理论剖析Transformer表达能力的边界,揭示大模型底层机制,值得算法研究者细读。
arXiv:2608.12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) i…
一眼看懂Transformer内部机制,用交互可视化亲手演示LLM的推理过程
arXiv:2408.04619v2 Announce Type: replace-cross Abstract: The Transformer architecture underpins modern large language models powering state-of-the-ar…
Transformer诞生九年后,一批初创公司正押注LLM的下一场革命,看懂前沿风向就看这篇。
MIT Technology Review’s What’s Next series looks across industries, trends, and technologies to give you a first look at the future. You can read the …
用马尔可夫链给Transformer层做路由,动态跳过冗余层,推理效率有望大幅提升。
arXiv:2608.05872v1 Announce Type: cross Abstract: Standard Large Language Models (LLMs) execute layers sequentially. Dynamic layer routing, i.e. searc…
百万参数小模型也能用思维链推理?这项研究让推理机制分析不再只属于大模型。
arXiv:2608.04980v1 Announce Type: cross Abstract: We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we cal…
MoE训练新方案:Sinkhorn梯度下降替代AdamW,减少优化器状态内存,让大模型训练更省显存。
arXiv:2608.04407v1 Announce Type: cross Abstract: Memory-efficient matrix optimizers such as Sinkhorn gradient descent remove most AdamW optimizer sta…
从词元到混合专家模型,一张按依赖顺序梳理的LLM术语地图,适合想系统搞懂大模型核心概念的读者。
TL;DR — A glossary to actually understand the terms you hit when reading about LLMs: token, embedding, attention, KV cache, GQA, MoE, quantization and…
把 Transformer 参数与激活约束到超球面,nGPT 带来表示学习新范式,一文看懂归一化训练的关键设计。
arXiv:2608.01284v1 Announce Type: cross Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining mo…
谷歌曾抢先一年做出ChatGPT却不敢发布?Transformer作者亲述内幕,看AI巨头如何错失先机。
终究还是喜提了「美国豆包」
医学影像融合迎来意图驱动新思路:扩散Transformer多模态网络,ACM MM 2026收录,技术细节扎实,值得算法研究者细读。
arXiv:2607.28565v1 Announce Type: new Abstract: Medical image fusion aims to integrate complementary information from diverse imaging modalities to su…
在模拟城市中边玩边学,轻松理解LLM内部运作机制!
Article URL: https://laurentiugabriel.github.io/token-town/ Comments URL: https://news.ycombinator.com/item?id=49068477 Points: 10 # Comments: 4
揭秘Transformer大模型如何像人类一样解决算术问题,并展示用人类策略进行模型调优的路径。
arXiv:2607.17166v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) continue to achieve state-of-the-art performance across…
Transformer新突破:链式计算与结构化上下文窗口,让AI规划能力再升级!
arXiv:2607.17710v1 Announce Type: new Abstract: Large Language Models (LLMs) have had a remarkable impact across many areas of machine learning. Howev…
提出关联递归记忆方法,让大模型上下文窗口突破长度限制,高效处理超长序列
arXiv:2607.11614v1 Announce Type: cross Abstract: Extending the context length of large language models (LLMs) is critical for many real-world applica…
NVIDIA官方详解硬件感知的大模型设计,平衡吞吐量与延迟的Pareto前沿策略。
Article URL: https://developer.nvidia.com/blog/ai-model-co-design-hardware-friendly-llm-design/ Comments URL: https://news.ycombinator.com/item?id=488…
为AI模型装上可控“开关”:新方法按类别精准限制危险知识,不牺牲正常能力,给安全对齐带来新思路。
Article URL: https://www.anthropic.com/research/off-switch-dual-use Comments URL: https://news.ycombinator.com/item?id=48841308 Points: 2 # Comments: …
一项新研究揭示Transformer在训练中会收敛到不变的算法核心,为理解模型行为与电路机制提供新视角。
arXiv:2602.22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the s…
Transformer融合谐波特征,实时精准定位磁粒子成像导管,介入手术导航新突破。
arXiv:2607.02919v1 Announce Type: cross Abstract: Magnetic particle imaging (MPI) enables real-time, radiation-free tracking of magnetic nanoparticle-…