1
The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data
斯坦福团队构建高质量金融披露数据集,兼顾布局忠实与token高效,为NLP预训练提供新基准。
arXiv:2606.18192v1 Announce Type: new Abstract: As high-quality public web corpora become increasingly exhausted, clean long-context documents have be…