High-Layer Attention Pruning with Rescaling
大模型剪枝新招:高层注意力剪枝加缩放,推理提速不损精度,论文原文速看。
arXiv:2507.01900v3 Announce Type: replace-cross Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), signifi…
大模型剪枝新招:高层注意力剪枝加缩放,推理提速不损精度,论文原文速看。
arXiv:2507.01900v3 Announce Type: replace-cross Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), signifi…
揭秘大模型性能增长的底层规律,一文读懂 Scaling Laws 如何指导 LLM 训练与架构设计。
Article URL: https://aidoses.substack.com/p/scaling-laws-the-law-behind-every Comments URL: https://news.ycombinator.com/item?id=49111433 Points: 2 # …
揭示大模型优势源于约束引导推理,而非单纯规模扩大,颠覆传统认知。
arXiv:2606.26108v1 Announce Type: new Abstract: Larger language models consistently outperform smaller ones on reasoning benchmarks, yet the reasoning…
6G网络与轻量语言模型结合,探索极小型AI推理极限,通信智能化必读研究。
arXiv:2603.02156v2 Announce Type: replace-cross Abstract: Emerging 6G visions, reflected in ongoing standardization efforts within 3GPP, IETF, ETSI, I…
提出ReM-MoA机制,用推理记忆解决混合代理规模化难题,提升多智能体协作效率。
arXiv:2606.24437v1 Announce Type: new Abstract: Mixture-of-Agents (MoA) architectures improve inference-time scaling by organizing multiple LLM agents…
颠覆深度缩放传统认知:层间相似性导致“越深越差”的逆缩放现象,ICML 2026最新发现。
arXiv:2602.05970v2 Announce Type: replace Abstract: Neural scaling laws relate loss to model size in large language models (LLMs), yet depth and width…