🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Data · Evaluation

The Stanford EDGAR Filings Dataset

Free while signed in. Answers cite the passages they came from.

First page
The Stanford EDGAR Filings Dataset
The curator’s take

Clean, long-context documents remain scarce for pretraining, especially in finance. This release reconstructs U.S. SEC corporate and financial disclosures into layout-faithful, token-efficient MultiMarkdown, publishing 152B tokens in SEFD-v1 out of an estimated 550B-token archive spanning 18.5M filings, with less than 0.1% overlap with Common Crawl corpora. It also ships two derived benchmarks, EDGAR-Forecast for numerical forecasting and EDGAR-OCR for financial table transcription, to support financial reasoning, forecasting, and document understanding.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack