Training nGPT
把 Transformer 参数与激活约束到超球面,nGPT 带来表示学习新范式,一文看懂归一化训练的关键设计。
arXiv:2608.01284v1 Announce Type: cross Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining mo…
把 Transformer 参数与激活约束到超球面,nGPT 带来表示学习新范式,一文看懂归一化训练的关键设计。
arXiv:2608.01284v1 Announce Type: cross Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining mo…
抖音多模态嵌入模型技术报告,揭秘工业级搜索推荐背后的向量表示学习。
arXiv:2608.02148v1 Announce Type: cross Abstract: Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and…
利用冻结大语言模型生成可复用身份表示,实现精准且高效的实体对齐,为知识图谱融合提供新思路
arXiv:2607.25579v1 Announce Type: cross Abstract: Entity alignment (EA) identifies entities across knowledge graphs (KGs) that refer to the same real-…
新型越狱攻击方法,通过混合有害和无害的隐藏状态绕过安全对齐,揭示大模型防御漏洞。
arXiv:2508.10029v3 Announce Type: replace-cross Abstract: Safety-aligned large language models can still be manipulated through white-box intervention…
颠覆认知:GNN本质是高级启发式算法而非特征学习器,理论推导+实验验证给出新视角
arXiv:2601.13465v4 Announce Type: replace Abstract: Graph neural networks are usually treated as auxiliaries for combinatorial optimization: they imit…
颠覆知识蒸馏常规认知:预训练表征只有等价类意义,匹配坐标是伪命题
arXiv:2607.03572v1 Announce Type: cross Abstract: Knowledge distillation is usually framed as a choice of what to match in the teacher - its logits, h…
联邦微调藏风险!图表示学习竟能放大模型操纵攻击,LLM安全防线告急。前沿研究不可错过。
arXiv:2605.07961v2 Announce Type: replace Abstract: Federated fine-tuning (FFT) has emerged as a privacy-preserving paradigm for collaboratively adapt…
把纵向医学影像建模成图,用 AI 预判治疗反应,为个性化医疗带来全新可能。
arXiv:2607.04912v1 Announce Type: cross Abstract: In patients with breast cancer, pathological complete response (pCR) has been established as a clini…
告别一刀切融合,提出针对时间事件建模的跨模态表示对齐新范式,精准提升多模态生存预测性能。
arXiv:2606.15038v1 Announce Type: new Abstract: Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modal…
自我监督概念发现新方法,用偏好学习破解可解释性与可扩展性的两难困境。
arXiv:2606.14586v1 Announce Type: new Abstract: Current representation learning paradigms force a fundamental compromise: self-supervised methods scal…
ICML 2026论文提出潜空间规划的时间直化方法,有效提升长期规划的连贯性与效率。
arXiv:2603.12231v2 Announce Type: replace Abstract: Learning good representations is essential for latent planning with world models. While pretrained…
提出通过类别-姿态分解实现等变表示学习,在变换不变性与泛化性上取得新突破。
arXiv:2207.03116v4 Announce Type: replace Abstract: We introduce a general method for learning representations that are equivariant to symmetries of d…
无需训练,仅通过反转输入文本,就能显著提升解码器LLM的文本嵌入质量。
arXiv:2606.05858v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have opened new avenues for generating training-free t…
医学视觉问答新突破:噪声感知学习提升视觉表示鲁棒性
arXiv:2606.05535v1 Announce Type: cross Abstract: Medical visual question answering (Med-VQA) has strong potential for clinical decision support by en…
自监督学习新突破:CoralBay开创CT基础模型,无需标注数据即可高效训练医学影像AI
arXiv:2606.03888v1 Announce Type: new Abstract: Self-supervised learning has enabled large-scale pre-training on 2D natural images, producing general-…
基于码本的连续用户表示方法,让LLM实现高效个性化生成,无需微调模型参数。
arXiv:2602.00742v2 Announce Type: replace Abstract: User modeling characterizes individuals through their preferences and behavioral patterns to enabl…
针对任意图结构学习表示的通用框架,原创性高,已被ICML 2026接收。
arXiv:2512.11561v2 Announce Type: replace Abstract: Generalizing pretrained models to unseen datasets without retraining is a central challenge toward…
KDD 2026接收,提出用最小充分表示学习为LLM高效合成领域数据,降低人工成本。
arXiv:2605.30039v1 Announce Type: new Abstract: Large Language Models have demonstrated remarkable progress in general-purpose capabilities and can ac…
超越思维链?新研究提出"重写"范式,将生成式多模态嵌入统一为通用接口,重塑AI推理与表示方式。
arXiv:2604.22280v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have emerged as a promising foundation for universal mult…
提出一种无需训练的向量量化新方法,利用高斯VAE实现高效表示学习,为量化领域开辟新路径。
arXiv:2512.06609v3 Announce Type: replace Abstract: Vector-quantized variational autoencoders (VQ-VAEs) are discrete autoencoders that compress images…