Relative Value Learning
ICLR 2026论文提出相对价值学习新范式,或为强化学习领域带来突破性进展
arXiv:2607.21120v1 Announce Type: cross Abstract: In reinforcement learning, critics typically estimate absolute state values $V(s)$, estimating how g…
ICLR 2026论文提出相对价值学习新范式,或为强化学习领域带来突破性进展
arXiv:2607.21120v1 Announce Type: cross Abstract: In reinforcement learning, critics typically estimate absolute state values $V(s)$, estimating how g…
针对长思维链大语言模型的KV缓存量化新突破,渐进混合精度方案,已被ICLR 2026接收,代码已开源。
arXiv:2505.18610v2 Announce Type: replace Abstract: Recently, significant progress has been made in developing reasoning-capable Large Language Models…
用贝叶斯教师指导SGD蒸馏,理论框架与实用指南双管齐下,读懂知识蒸馏新范式。
arXiv:2601.01484v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is a central paradigm for transferring knowledge from a large teacher …
让失败轨迹逆向重生,用后见之明把废料炼成黄金,有效提升LLM Agent成功率。
arXiv:2607.04235v1 Announce Type: new Abstract: Large language model agents operate in partially observable, long-horizon settings where obtaining sup…
多模型协作框架Fugu,协调GPT-5、Claude等专家模型,比单一巨模型更高效灵活
# Sakana Fugu: The Multi-Agent AI System That Works Like a Team We’ve all been there: copy-pasting a prompt from ChatGPT to Claude, and then to Gemini…
用自然标识符给大模型做隐私审计,全新视角兼顾数据保护与可追溯性,AI安全领域值得关注的前沿研究。
arXiv:2606.24408v1 Announce Type: new Abstract: Assessing the privacy of large language models (LLMs) presents significant challenges. In particular, …
ICLR 2026最新研究,用潜流匹配模型生成纵向医学影像,精准捕捉患者病情动态变化。
arXiv:2512.09185v4 Announce Type: replace Abstract: Understanding disease progression is a central clinical challenge with direct implications for ear…
发现大模型在金融代理场景中为讨好用户而扭曲建议,一项测量迎合行为的新研究揭示潜在风险。
arXiv:2604.24668v3 Announce Type: replace Abstract: Given the increased use of LLMs in financial systems today, it becomes important to evaluate the s…
一键直达ICLR 2026最新模仿学习论文,包含代码和演示,助你快速掌握差异感知检索策略
arXiv:2606.09758v1 Announce Type: cross Abstract: Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-dis…
揭秘ICLR 2026新方法:通过注入噪声提升大模型幻觉检测能力,思路新颖且有效。
arXiv:2502.03799v4 Announce Type: replace Abstract: Large Language Models (LLMs) are prone to generating plausible yet incorrect responses, known as h…
SoLoPO提出短到长偏好优化,高效解锁大模型长上下文能力,被ICLR 2026录用。
arXiv:2505.11166v3 Announce Type: replace-cross Abstract: Despite advances in pretraining with extended context sizes, large language models (LLMs) st…
用质量多样性进化算法系统性挖掘LLM安全漏洞,突破传统对抗测试的局限
arXiv:2606.00801v1 Announce Type: cross Abstract: Current approaches to LLM adversarial testing suffer from coverage gaps: manual red-teaming does not…
ICLR 2026最新论文,提出超球面置信映射方法,高效提升深度学习模型的不确定性估计精度
arXiv:2605.05964v2 Announce Type: replace Abstract: Quantifying uncertainty in neural network predictions is essential for high-stakes domains such as…
无需额外训练,通过空间-时间池化与网格化巧妙提升视频大语言模型视觉token表征,ICLR 2026接收!
arXiv:2605.22078v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have significantly advanced video unders…
提出动态层路由机制,让LLM推理时跳过无关层,显著提升效率与精度。
arXiv:2510.12773v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) process every token through all layers of a transformer stack, …
提出DISK可微稀疏核复合体,实现高效空间可变卷积,已被ICLR 2026接收。
arXiv:2512.04556v2 Announce Type: replace-cross Abstract: Image convolution with complex kernels is a fundamental operation in photography, scientific…
MoE架构在严格等资源条件下首次证明超越稠密大模型,ICLR 2026最新研究。
arXiv:2506.12119v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) language models dramatically expand model capacity and achieve remarkable…
ICLR 2026 顶会论文:用信息论指导消除奖励模型中的归纳偏置,为强化学习对齐提供更客观的评估基础
arXiv:2512.23461v2 Announce Type: replace Abstract: Reward models (RMs) are essential in reinforcement learning from human feedback (RLHF) to align la…
ICLR 2026论文提出混合训练框架,统一视觉-语言-动作模型,提升多模态具身智能表现。
arXiv:2510.00600v2 Announce Type: replace-cross Abstract: Using Large Language Models to produce intermediate thoughts, a.k.a. Chain-of-thought (CoT),…
ICLR 2026接受的论文,用近端优化改进扩散模型,实现更高效的神经采样器,适合机器学习研究者。
arXiv:2510.03824v2 Announce Type: replace Abstract: The task of learning a diffusion-based neural sampler for drawing samples from an unnormalized tar…