CommonCanvas
Free while signed in. Answers cite the passages they came from.

Releases CommonCanvas, a text-to-image dataset composed entirely of Creative-Commons-licensed images.
CC-only training data: Every image is Creative Commons-licensed, providing a clean-license dataset for commercial and research T2I training.
Scale despite licensing constraints: Curates hundreds of millions of images despite the CC-only constraint, dispelling the myth that legal T2I training requires permissive copyrighted data.
Strong baseline models: Trains SD-style models on CommonCanvas that reach competitive quality, demonstrating CC data is sufficient for SoTA T2I.
Policy contribution: Provides a practical counterexample to the argument that copyrighted training data is necessary - important as copyright litigation reshaped the AI-data landscape.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack