🚀NEW LABGetting Started with Claude AgentsStart lab
Data · Evaluation

The Stanford EDGAR Filings Dataset

First page
The Stanford EDGAR Filings Dataset
Paper summary

Clean, long-context documents remain scarce for pretraining, especially in finance. This release reconstructs U.S. SEC corporate and financial disclosures into layout-faithful, token-efficient MultiMarkdown, publishing 152B tokens in SEFD-v1 out of an estimated 550B-token archive spanning 18.5M filings, with less than 0.1% overlap with Common Crawl corpora. It also ships two derived benchmarks, EDGAR-Forecast for numerical forecasting and EDGAR-OCR for financial table transcription, to support financial reasoning, forecasting, and document understanding.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack