AI data startup Micro1 reaches $500M gross run rate amid AI training boom
AI数据公司Micro1营收狂奔至5亿美元,揭示合成数据与训练需求爆发。
Surging demand for AI training data is driving rapid growth for the startup and its rivals.
AI数据公司Micro1营收狂奔至5亿美元,揭示合成数据与训练需求爆发。
Surging demand for AI training data is driving rapid growth for the startup and its rivals.
图灵奖得主萨顿泼冷水:合成数据模拟不了真实世界的复杂,AI巨头该冷静了。
IT之家 8 月 19 日消息,里奇 · 萨顿(Rich Sutton)是当今人工智能热潮背后关键技术的先驱之一,但如今,他认为科技巨头为了让 AI 持续发展而采取的路线存在严重问题。 这位加拿大计算机科学家、图灵奖得主在当地时间周二发布的红杉资本播客节目中表示,AI 行业越来越依赖合成训练数据,这…
用LLM合成训练数据,破解自动修Bug的数据瓶颈,思路新颖且实战价值高。
arXiv:2505.07372v3 Announce Type: replace-cross Abstract: This paper presents a novel methodology for enhancing Automated Program Repair (APR) through…
针对多语言提示下大模型表现不一的难题,用定向合成数据提升模型智能,值得关注。
arXiv:2608.15964v1 Announce Type: cross Abstract: Language-specific competency (LSC) is the phenomenon of a language model performing better or worse …
用LLM智能体自动发现列间约束,让合成表格数据更真实可信,为数据增强与隐私保护提供新思路。
arXiv:2608.15109v1 Announce Type: new Abstract: Generating structurally valid synthetic tabular data remains difficult: outputs with high statistical …
用AI合成患者数据破解小样本难题,精准还原妊娠凝血动态变化,医学研究新利器。
arXiv:2604.07557v2 Announce Type: replace Abstract: Small longitudinal cohorts, common in maternal health, rare diseases, and early-phase trials, limi…
临床通信遇上AI:用大模型合成数据训练,破解医疗文本处理的数据难题,技术路径与实战案例一次讲透。
arXiv:2608.05993v1 Announce Type: new Abstract: Much clinical value is conveyed not through structured records but through communication: exchanges in…
用合成数据驯服农业领域多语言大模型,破解低资源语言问答难题,实验严谨、方法可复用。
arXiv:2507.16974v3 Announce Type: replace-cross Abstract: Enabling farmers to access accurate agriculture-related information in their native language…
通过合成指令数据扩展预训练规模,突破传统监督训练数据瓶颈的新方法
arXiv:2601.22146v2 Announce Type: replace-cross Abstract: Due to limited supervised training data, large language models (LLMs) are typically pre-trai…
触觉手套(图源/企业) 本文约 2400 字,建议阅读 6 分钟 作者丨欧雪 编辑丨袁斯来 硬氪获悉,空间智能公司「大衍科技」近期完成数千万元天使轮融资,由松禾资本领投,浙江省省金控与广州番禺创新基金等国资背景机构参与。资金将主要用于触觉大模型研发、机器人数据产线建设及团队扩张。 「大衍科技」202…
利用大语言模型低成本生成调查数据时,准确性因问题波动,本文研究如何最优分配固定的人类样本预算以提升估计效果。
arXiv:2604.17267v2 Announce Type: replace Abstract: Large Language Models can generate synthetic survey responses at low cost, but their accuracy vari…
用结构化提示提升合成日语咨询对话质量,并探索自动评估方法,为对话系统研究提供新思路。
arXiv:2507.02950v3 Announce Type: replace-cross Abstract: Large language models (LLMs) may support counseling training, yet evidence from Japanese-lan…
针对智能体任务失败根因,提出按能力缺口定向训练的新范式,为Agent训练提供精准解法。
arXiv:2604.05336v2 Announce Type: replace Abstract: Models often fail to complete agentic tasks because they lack core capabilities required by the ta…
揭秘智能体信息检索新范式:将数据驱动优化从模型转向工作环境本身
arXiv:2607.00016v1 Announce Type: cross Abstract: Information localization within massive repositories is a cornerstone of agentic LLM systems. While …
揭示语言模型自我生成QA训练的隐藏脆弱性,一篇值得关注的AI研究论文。
arXiv:2606.32002v1 Announce Type: new Abstract: Language models are increasingly taught from synthetic question--answer (QA) supervision: a model gene…
用合成交互数据破解LLM个性化扩展难题,数据瓶颈的新解法。
arXiv:2602.12394v2 Announce Type: replace Abstract: Personalized prompting offers large opportunities for deploying large language models (LLMs) to di…
无需人工标注,DocArena把原始文档变成可控训练场,让文档搜索代理练出更强检索能力,值得关注。
arXiv:2606.26122v1 Announce Type: new Abstract: Recent methods train search agents via reinforcement learning from (question, answer, evidence) tuples…
开源工具助你高效清洗与生成合成数据,优化大模型后训练流程。
arXiv:2606.21631v1 Announce Type: cross Abstract: Data curation is a critical part of post-training pipelines for large language models, yet existing …
评估合成数据在胎儿脑MRI分割域泛化中的效果,探索提升模型跨域鲁棒性的新思路。
arXiv:2411.06842v3 Announce Type: replace-cross Abstract: Fetal brain tissue segmentation from magnetic resonance imaging (MRI) is crucial for studyin…
用AI生成逼真医疗对话与病历对,解决数据隐私难题,医学NLP新突破。
arXiv:2508.01401v2 Announce Type: replace-cross Abstract: Physicians spend significant time documenting clinical encounters, a burden that contributes…