1
How much of an LLM-generated clinical corpus is actually new? A production-scale measurement of content redundancy for provenance classification
临床语料里有多少是AI抄来的?实测大模型生成内容的冗余度,给数据来源打假。
arXiv:2606.29605v1 Announce Type: new Abstract: Clinical machine learning increasingly relies on training corpora generated by large language models (…