Is the future of data centers portable? Runware builds a pod to find out
便携式AI数据中心,让算力随需部署,大幅缩短建设周期,适合快速扩展的AI团队。
On Tuesday, AI infrastructure company Runware announced the launch of its own modular data center called Sonic Inference Pod.
便携式AI数据中心,让算力随需部署,大幅缩短建设周期,适合快速扩展的AI团队。
On Tuesday, AI infrastructure company Runware announced the launch of its own modular data center called Sonic Inference Pod.
亿级日活App靠跨云架构砍掉75% GPU集群,AI出海破解算力成本倒挂的真实解法。
出海AI,正被“三重算力锁链"捆死
用无分支控制流重写NCCL路由,把大规模集群同步抖动压向物理0ns,AI网络架构的硬核蓝图不容错过。
Article URL: https://github.com/PJHkorea/branchless-nccl-router Comments URL: https://news.ycombinator.com/item?id=48810692 Points: 2 # Comments: 1
Netris获a16z 1500万美元A轮融资,解决GPU集群每日链路配置痛点,加速AI云部署
Netris provides software that runs on network switches, and offers a platform that helps neocloud operators reduce the time it takes to go live.
欧洲史上最大AI超算启动,35台英伟达HPC系统惠及300万+研究人员,加速生成式AI、气候建模与医疗等关键领域
IT之家 6 月 22 日消息,英伟达今晚宣布,欧洲创纪录的 35 台英伟达 AI HPC 超级计算机正式启动建设。建成后,超过 300 万名研究人员将获得下一代算力基础设施,用于推动全欧洲 AI 发展、加快科学研究和促进工业创新。 35 台系统构成欧洲有史以来规模最大的一次年度超算扩建,覆盖国家级…
从SRE视角剖析AI推理基础设施,揭秘真实工作流与技能树,适合想转型该领域的工程师
Hi, I currently work on a GenAI platform for one of the largest local industrial companies. My daily work mostly involves building inference infrastru…
真实生产环境下的LLM预训练运维经验,504块GPU集群从故障检测到恢复的实证分析。
arXiv:2605.09370v2 Announce Type: replace-cross Abstract: Large-scale AI training is now fundamentally a distributed systems problem, and hardware fai…
面向共享GPU集群,提出连续自适应方法优化大模型服务SLO,降低延迟与成本
arXiv:2604.16400v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) are increasingly adopted in edge intelligence to power domai…
OpenAI官方分享训练大型神经网络的核心技术与工程挑战,GPU集群同步计算的关键方法。
Large neural networks are at the core of many recent advances in AI, but training them is a difficult engineering and research challenge which require…