Show HN: Data enrichment with AI for Pandas DataFrame
为Pandas DataFrame注入AI能力,自动完成数据增强与特征工程,数据预处理从此更高效。
Article URL: https://github.com/mljar/enrichment Comments URL: https://news.ycombinator.com/item?id=49372451 Points: 1 # Comments: 0
为Pandas DataFrame注入AI能力,自动完成数据增强与特征工程,数据预处理从此更高效。
Article URL: https://github.com/mljar/enrichment Comments URL: https://news.ycombinator.com/item?id=49372451 Points: 1 # Comments: 0
针对多语言提示下大模型表现不一的难题,用定向合成数据提升模型智能,值得关注。
arXiv:2608.15964v1 Announce Type: cross Abstract: Language-specific competency (LSC) is the phenomenon of a language model performing better or worse …
用几何过滤筛掉低质量LLM生成样本,少样本分类精度立竿见影。
arXiv:2608.13866v1 Announce Type: new Abstract: Large language models (LLMs) can generate synthetic training data for text classification, but the qua…
随机权重平均遇上数据增强,简单组合就能显著提升模型泛化,炼丹必备技巧。
arXiv:2608.14373v1 Announce Type: new Abstract: The symmetries of a learning task have become an important factor in designing modern deep learning so…
开源大模型以细粒度辩论实现可信数据扩充,为心理健康与网络内容安全标注开辟新路径。
arXiv:2512.06227v3 Announce Type: replace-cross Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) appli…
把物理规律嵌入自监督增强,让科学成像模型学到的特征更可解释、更稳健,是AI4Science方向值得关注的新思路。
arXiv:2607.28868v1 Announce Type: new Abstract: Data augmentations define the invariances learned by self-supervised learning (SSL). Standard augmenta…
用对抗不确定性的方式引导LLM生成语义增强数据,大幅提升异质处理效应估计精度
arXiv:2607.26599v1 Announce Type: new Abstract: Estimating heterogeneous treatment effects is central to targeted interventions, such as personalized …
比较研究如何通过数据增强提升大语言模型对电子病历关键医学术语的识别与排序能力
arXiv:2502.16022v3 Announce Type: replace Abstract: OpenNotes gives patients access to their EHR notes, but dense medical jargon limits comprehension.…
用微调大模型生成反事实解释,为健康干预设计和数据增强提供新思路,实用性强。
arXiv:2601.14590v3 Announce Type: replace Abstract: Counterfactual explanations (CFEs) provide human-centric interpretability by identifying the minim…
不增加昂贵教师模型推理,用少量教师步数实现高效on-policy数据增强,让智能体后训练更省钱更精准。
arXiv:2607.04574v1 Announce Type: cross Abstract: For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about whi…
利用CNN训练中的不确定性作为数据增强新策略,为提升模型泛化能力提供创新思路。
arXiv:2509.05238v2 Announce Type: replace-cross Abstract: Deep learning (DL) has transformed neuroimaging by delivering state-of-the-art performance w…
针对真实临床数据长尾分布,引入生成式数据增强与标签共现建模,提升胸片多标签分类性能。
arXiv:2607.00975v1 Announce Type: cross Abstract: Chest X-ray multi-label classification is a core task in intelligent medical imaging diagnosis. Howe…
贝叶斯神经网络与等变性结合,数据增强理论新突破,提升模型鲁棒性
arXiv:2606.26273v1 Announce Type: new Abstract: Symmetries are important for many deep learning tasks, ranging from applications in the sciences to me…
用AI生成逼真医疗对话与病历对,解决数据隐私难题,医学NLP新突破。
arXiv:2508.01401v2 Announce Type: replace-cross Abstract: Physicians spend significant time documenting clinical encounters, a burden that contributes…
论文提出前缀式数据增强方法扩展nnU-Net,显著提升医学图像分割精度与泛化能力。
arXiv:2606.10713v1 Announce Type: cross Abstract: The nnU-Net has demonstrated continuous success in medical segmentation tasks, which heavily rely on…
提出训练后的数据增强不变性方法,提升模型泛化能力,不依赖额外训练成本。
arXiv:2505.11702v3 Announce Type: replace Abstract: This work develops a framework for post-training augmentation invariance, in which our goal is to …
基于扩散模型实现几何感知的表格数据合成,为隐私保护下的数据共享提供高效方案
arXiv:2606.02607v1 Announce Type: cross Abstract: Tabular synthesis is critical for privacy-preserving sharing and augmentation, yet diffusion models …
利用跨域事件生成合成数据,为大规模推荐系统提供低成本、高效率的训练方案。
arXiv:2606.00282v1 Announce Type: cross Abstract: Large-scale recommendation systems operate across diverse domains, yet they face the challenges of d…
ICML 2026 论文提出基于 Foundation VAE 的统一框架,实现 3D CT 图像高质量重建、数据增强与生成,将推动医学影像 AI 发展。
arXiv:2605.30893v1 Announce Type: new Abstract: Variational autoencoders (VAEs) compress high resolution CT volumes into compact latents while preserv…
用稀疏自编码器在LLM特征空间合成多样化数据,实现“少即是多”的高效数据扩充新方法。
arXiv:2602.10388v3 Announce Type: replace-cross Abstract: The diversity of post-training data is critical for effective downstream performance in larg…