Douyin Multimodal Embedding Model Technical Report
抖音多模态嵌入模型技术报告,揭秘工业级搜索推荐背后的向量表示学习。
arXiv:2608.02148v1 Announce Type: cross Abstract: Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and…
抖音多模态嵌入模型技术报告,揭秘工业级搜索推荐背后的向量表示学习。
arXiv:2608.02148v1 Announce Type: cross Abstract: Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and…
OpenAI将发布Hugging Face被黑事件技术报告,揭秘AI平台安全漏洞细节。
OpenAI:我们注意到,围绕抱抱脸系统(Hugging Face)被入侵事件流传着许多疑问和各种猜测。我们正与外部顾问合作,并在安全与安保委员会的监督下,继续进行全面审查。审查完成后,我们计划在未来几周内发布一份技术报告,分享我们从中汲取的经验与见解。(新浪财经)
子十亿参数多模态AI代理Octopus v3技术报告,揭秘如何在设备端实现高效推理。
arXiv:2404.11459v3 Announce Type: replace Abstract: A multimodal AI agent is characterized by its ability to process and learn from various types of d…
LLM驱动GPU内核自动生成,工程优化是关键——比赛级技术报告揭示提示工程与循环生成创新
arXiv:2607.17979v1 Announce Type: cross Abstract: Large language models (LLMs) can assist GPU kernel generation, but their practical effectiveness dep…
欧洲主权大模型Soofi仅用2个月完成训练,技术报告详解其高效架构与开源细节。
Article URL: https://huggingface.co/spaces/Soofi-Project/Pretraining-Tech-Report Comments URL: https://news.ycombinator.com/item?id=48870978 Points: 8…
面向AI代理的原生技能生产平台SkillFab,让技能制作更智能高效。
arXiv:2607.03780v1 Announce Type: cross Abstract: SkillFab is an agent-native platform for turning missing capabilities into reviewed, reusable Agent …
小米发布新一代GUI智能体,解锁跨应用真实操作!技术报告详解模型能力,必读干货。
arXiv:2606.31410v1 Announce Type: new Abstract: Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-en…
Sakana AI最新技术报告,揭秘河豚(Fugu)模型的创新设计与研究突破。
arXiv:2606.21228v1 Announce Type: new Abstract: The capabilities of frontier Large Language Models (LLMs) continue to advance, with different provider…
最新技术报告《NeuroClaw》公开,深入探讨该系统的核心设计与实现细节。
arXiv:2604.24696v3 Announce Type: replace Abstract: Agentic artificial intelligence systems promise to accelerate scientific workflows, but neuroimagi…
访问arXiv论文库中的最新技术报告,快速获取扩散模型从左到右生成的前沿研究,学术必备
arXiv:2606.11552v1 Announce Type: cross Abstract: Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their …
通用提示改进可能适得其反,这篇论文用严谨的评估驱动迭代方法,给出LLM应用避坑指南。
arXiv:2601.22025v2 Announce Type: replace-cross Abstract: Evaluating Large Language Model (LLM) applications differs from conventional software testin…
快手发布Keye-VL-2.0多模态大模型技术报告,31页详解架构与训练细节
arXiv:2606.10651v1 Announce Type: new Abstract: We introduce Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation m…
MOSS大模型音频能力技术报告,揭秘音频理解与生成新突破
arXiv:2606.01802v1 Announce Type: cross Abstract: MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understandin…
CASTLE2026比赛团队技术报告,详解WDL方法设计思路与实验细节,为竞赛方案提供参考。
arXiv:2606.00712v1 Announce Type: new Abstract: The CASTLE Challenge @ EgoVis 2026 evaluates long-form egocentric video question answering over 600+ h…
百万级LLM训练与服务的托管基础设施技术报告,揭秘超大规模模型管理的核心架构。
arXiv:2605.13779v2 Announce Type: replace-cross Abstract: We present MindLab Toolkit (MinT), a managed infrastructure system for Low-Rank Adaptation (…
韩国团队Raon-Speech语音技术报告,揭示多模态语音AI前沿突破
arXiv:2605.23912v1 Announce Type: cross Abstract: We present Raon-Speech, a top-performing 9B-parameter speech language model (SpeechLM) for English a…
通过可验证任务缩放的新方法,显著提升大语言模型推理能力,来自InternBootcamp的技术报告亮点频出。
arXiv:2508.08636v2 Announce Type: replace Abstract: Large language models (LLMs) have revolutionized artificial intelligence by enabling complex reaso…
通义千问发布DeepResearch技术报告,揭秘面向长周期深度搜索的端到端agentic大模型训练框架
arXiv:2510.24701v3 Announce Type: replace-cross Abstract: We present Tongyi DeepResearch, an agentic large language model, which is specifically desig…
EgoVis 2026 CASTLE挑战赛亚军方案的技术报告,详解多模态场景理解新方法MARS,视觉AI进阶必读
arXiv:2605.18176v1 Announce Type: new Abstract: This report presents MARS, short for Multimodal Agentic Reasoning with Source selection, our system fo…
GPT-5.3系统卡正式发布,详解最新模型能力、安全评估与技术细节