Crowd Workers Widely Use LLMs for Text Production
Free while signed in. Answers cite the passages they came from.

Empirical evidence that 33-46% of MTurk crowd workers used LLMs on text tasks.
LLM-generated contamination: Estimates that a third to almost half of crowd-worker text production involved LLMs - a massive data quality issue.
Benchmark contamination risk: Implications for NLP datasets produced via crowdsourcing, potentially invalidating many "human baseline" numbers.
Methodology: Uses statistical analysis comparing completion times, stylistic features, and output consistency to estimate LLM usage.
Community wake-up: Sparked widespread discussion about the future of human-generated data and the need for AI-usage detection.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack