InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers
一份针对地理分布式大模型数据中心的可持续规划方案,直击算力扩张与能耗平衡的痛点。
arXiv:2608.12915v1 Announce Type: cross Abstract: The rapid growth of LLM inference is shifting sustainability concerns from one-time training to cont…
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs
破解多模态大模型算力瓶颈,看ParVL如何并行扩展与弹性分配计算资源。
arXiv:2608.04010v1 Announce Type: cross Abstract: Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either mod…
SOLAR: AI-Powered Speed-of-Light Performance Analysis
用AI实现光速级性能分析,突破传统瓶颈,硬件优化新范式。
arXiv:2606.26383v1 Announce Type: cross Abstract: How fast could a deep-learning model run on target hardware, and how far is today's implementation f…
OpenAI 升级 ChatGPT 记忆系统:算力降至 1/5,瞄准过时与错误两大痛点
ChatGPT记忆系统大升级,算力降至1/5,解决记忆过时与错误两大痛点。
IT之家 6 月 5 日消息,OpenAI 公司昨日(6 月 4 日)宣布升级 ChatGPT 记忆功能,新系统基于 Dreaming V3 机制, 重点改善记忆过时、准确性和大规模服务能力 。 ChatGPT 记忆系统原本用于记住用户偏好和长期信息,从而减少每次对话都要重新说明背景的麻烦。 Cha…
Efficient Pre-Training with Token Superposition
提出Token叠加技术,颠覆预训练效率瓶颈,大幅降低算力需求,LLM训练优化必读。
arXiv:2605.06546v2 Announce Type: replace Abstract: Pre-training of Large Language Models is often prohibitively expensive and inefficient at scale, r…