Diffract: Spectral View of LLM Domain Adaptation
用谱视角剖析大模型领域适应,Diffract方法为持续预训练提供新洞察。
arXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language model…
用谱视角剖析大模型领域适应,Diffract方法为持续预训练提供新洞察。
arXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language model…
稀疏注意力中因果信息流被重新定义,Stem提出新视角,为长上下文模型效率与准确性兼得提供关键思路。
arXiv:2603.06274v2 Announce Type: replace-cross Abstract: The quadratic computational complexity of self-attention remains a fundamental bottleneck fo…
提出多任务GRPO方法,有效提升大模型跨任务推理的可靠性,已被ICML 2026收录。
arXiv:2602.05547v2 Announce Type: replace-cross Abstract: RL-based post-training with GRPO is widely used to improve large language models on individu…
用生物动力学知识先验解决数据稀缺难题,为生物过程建模开辟新路径——ICML 2026 AI for Science 论文
arXiv:2607.20539v1 Announce Type: cross Abstract: While deep learning has accelerated drug discovery, its impact on biomanufacturing has been consider…
ICML'26提出概念集中方法,通过忠实表示干预提升AI可解释性,为模型控制提供新思路。
arXiv:2505.18672v2 Announce Type: replace Abstract: Representation intervention aims to localize and modify the representations that encode the underl…
新方法通过奖励感知群体缩放优化进化策略,显著提升LLM微调效率,已被ICML 2026 workshop收录。
arXiv:2607.19408v1 Announce Type: new Abstract: Using Evolutionary Strategies (ES) for fine-tuning large language models is attractive because it is m…
提出可证明安全的高效聚类方案,为大规模LLM推理保驾护航。
arXiv:2607.19704v1 Announce Type: new Abstract: Scaling LLM-based applications to millions of users is bottlenecked by the inference cost and latency …
提出利用事后经验将结果转化为行动,显著提升长周期语言智能体训练效果,ICML 2026录用论文。
arXiv:2607.16257v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a widely adopted technique for improving large language model…
低秩预训练大模型遭遇不稳定困境?这项ICML 2026研究提出原生低秩LLM预训练的稳定化方法,兼顾效率与质量。
arXiv:2602.12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose signif…
ICML 2026重磅论文:进化策略替代强化学习,开创大模型微调新范式。
arXiv:2509.24372v3 Announce Type: replace Abstract: Fine-tuning large language models (LLMs) for downstream tasks is an essential stage of modern AI d…
专家训练时长如何影响LLM模型合并效果?这篇ICML 2026 workshop论文揭示关键发现。
arXiv:2607.11997v1 Announce Type: new Abstract: Multi-task model merging combines separately trained expert models into a single model that handles al…
揭秘大模型安全探针为何在最终token失效:早于最后一层的内部表征已暴露漏洞
arXiv:2605.12726v2 Announce Type: replace Abstract: Final-token safety probes monitor a single hidden state after prompt prefill, but jailbreak prompt…
ICML 2026论文深入解析模型引导技术中的强度量化问题,为AI可解释性提供新视角。
arXiv:2602.02712v2 Announce Type: replace Abstract: A popular approach to post-training control of large language models (LLMs) is the steering of int…
从神经层面拆解大模型「拍马屁」行为的内部机制,一篇被ICML 2026研讨会接收的机械可解释性突破。
arXiv:2607.07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement e…
强化学习后训练如何催生组合推理策略,来自ICML workshop的前沿实证。
arXiv:2607.07646v1 Announce Type: new Abstract: Does RL post-training merely amplify primitive skills already latent in a base model, or can it compos…
ICML 2026接收:为LLM水印引入功率校准,告别传统启发式调优。
arXiv:2607.05694v1 Announce Type: cross Abstract: Logit-based watermarking is a widely used mechanism for identifying LLM generated content, yet its e…
ICML 2026论文为LLM Agent在社交困境中的合作能力打造首个系统化基准,揭示维持机制的关键设计。
arXiv:2604.15267v2 Announce Type: replace-cross Abstract: It is increasingly important that LLM agents interact effectively and safely with other goal…
从学习理论视角解析模型崩溃,重放机制能否成为破局关键?ICML 2026 新作给出严谨答案。
arXiv:2603.11784v2 Announce Type: replace Abstract: As scaling laws push the training of frontier large language models (LLMs) toward ever-growing dat…
追踪开源聊天LLM十二个版本的信任度变化,揭示模型行为随时间漂移的潜在风险
arXiv:2607.02587v1 Announce Type: cross Abstract: Model cards quote trust-benchmark scores without recording when they were measured, and the same num…
统一框架稳定智能体强化学习,ICML 2026收录,直击训练不稳定痛点。
arXiv:2602.21534v3 Announce Type: replace Abstract: Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for trai…