1
Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?
探究大模型公开词表能否反推隐藏训练语料的token分布,直击数据隐私与安全痛点
arXiv:2608.10690v1 Announce Type: new Abstract: Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model we…