SelFusion: Self-distillation for Diffusion Language Models
扩散语言模型的自蒸馏新方法登上ACL 2026,生成效率与质量或迎来新突破。
arXiv:2608.22898v1 Announce Type: new Abstract: Diffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) larg…
扩散语言模型的自蒸馏新方法登上ACL 2026,生成效率与质量或迎来新突破。
arXiv:2608.22898v1 Announce Type: new Abstract: Diffusion language models (DLMs) alleviate the inherent latency bottleneck of autoregressive (AR) larg…
无需训练的开放世界对象放置新方法,想象搜索机制兼顾效率与表现,视觉生成必读。
arXiv:2608.21543v1 Announce Type: cross Abstract: Object placement is critical in image composition, requiring spatially and semantically coherent pos…
用Flow Matching生成模型攻克3D医学图像中细小管状结构分割难题,方法新颖且应用价值高。
arXiv:2608.19965v1 Announce Type: cross Abstract: Segmentation of curvilinear anatomical structures in 3D medical images remains challenging due to co…
智能体与程序协同设计,重塑3D创作新范式,值得技术控深读。
arXiv:2608.17975v1 Announce Type: cross Abstract: Programmatic representations provide a compelling paradigm for 3D content creation, enabling fine-gr…
用生成模型加速蒙特卡洛采样,解锁复杂概率分布推断的新思路,值得关注。
arXiv:2608.07648v1 Announce Type: cross Abstract: Sampling high-dimensional probability distributions is a central task in scientific computing, with …
阿里Wan3.0公测,单次生成30秒视频,支持多模态全能参考,从镜头升级到完整故事创作。
IT之家 8 月 6 日消息,“阿里云”公众号今天(6 日)晚间宣布,全新一代视频生成模型 Wan3.0 开启公测。 Wan3.0 在生成时长、万能创作、全能参考与真实感等维度全面升级,单次可生成 30 秒视频,完整表达创作意图;在文本、图片、音频、视频四种基础模态之外, 首次支持 doc、xls、…
这项研究将3D建模与空间推理解耦,为AI生成更精细的3D场景提供新思路。
arXiv:2608.05242v1 Announce Type: new Abstract: In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D …
全球首个千亿MoE视频生成模型开源,6B激活媲美第一梯队,生成成本低至五毛,开启视频生成新范式。
10秒1080P,成本只要5毛钱
提出一种参数化扩散桥的新方法,扩展生成模型的理论边界。
arXiv:2607.22719v1 Announce Type: new Abstract: Multiplicative Gamma noise is a signal-dependent degradation in coherent imaging; synthetic aperture r…
提出新型Logit-Coordinate生成模型,专为混合连续与分类变量的表格数据设计,有望提升表格数据生成质量。
arXiv:2607.23348v1 Announce Type: cross Abstract: Mixed continuous--categorical data pose a representation problem for continuous generative models. F…
小米开源具身智能统一生成模型,可通吃四类任务,推理速度提升83倍,开源爱好者不容错过。
IT之家 7 月 15 日消息,小米今日发布 Xiaomi-Robotics-U0—— 一个拥有 380 亿参数 的多模态自回归具身生成基础模型,是 具身领域首个“通吃”四类任务的统一生成模型 ,打通了机器人图片和视频数据的生成与编辑链路。 具身场景生成(Scene Generation)—— 模型…
无需额外训练,推理阶段即可实现跨主体的精准运动迁移,为视频生成与动画制作带来简洁高效的解决方案。
arXiv:2607.11644v1 Announce Type: new Abstract: This work explores the motion transfer from one video to another, which is crucial in animation for di…
用扩散模型生成随机图信号,搞定图数据建模与生成任务,学术党值得一看。
arXiv:2607.06833v1 Announce Type: new Abstract: Sampling stochastic signals supported on a graph underlies many graph machine learning tasks, includin…
6款主流AI视频生成工具深度横评,可灵3.0在“演员吃面”测试中展现惊人物理模拟能力,Seedance 2.0引领全民狂欢,新手也能3分钟上手。
2026年,AI视频生成赛道已迈入全面爆发的成熟竞速期,Seedance 2.0 的出圈更是让AI视频变成一场“全民狂欢”。 短短两年间,AI视频从最初几秒的碎片化模糊画面,到如今分钟级长视频的连贯叙事、真实物理世界的精准还原,AI视频工具的迭代速度远超预期,AI视频工具完成了从“能用”到“好用”再…
IT之家 7 月 8 日消息,Meta 于当地时间 7 月 7 日宣布推出其全新的 AI 图像生成模型 Muse Image。 该模型来自 Meta 旗下专注人工智能的 Meta 超级智能实验室,也是该实验室推出的首个图像生成模型。该模型现已通过 Meta AI 应用免费提供,并同步登陆 Insta…
Google Voice 推出付费订阅方案并调整 Android 设备备份策略等。 查看全文
聚焦扩散模型分数匹配gap的论文,提供理论收紧方法,助你优化生成质量
arXiv:2607.04442v1 Announce Type: cross Abstract: Diffusion models (DMs) are a state-of-the-art generative method to approximately sample from an unkn…
解析离散扩散模型学习机制的论文,助你深入理解生成模型原理
arXiv:2607.05381v1 Announce Type: cross Abstract: What does a discrete diffusion model learn: a denoiser, a score ratio, or a bridge plug-in predictor…
用数据高效感知研究破解生成模型必要性,视觉AI进阶必读论文
arXiv:2512.08854v3 Announce Type: replace-cross Abstract: It has been hypothesized that achieving the data efficiency of human visual perception requi…
提出一种最优混合专家模型平均方法,提升条件生成模型的性能和稳定性,理论扎实。
arXiv:2607.04360v1 Announce Type: cross Abstract: Conditional generative models have emerged as powerful tools for sampling from target conditional di…