LLMs on University-Level Physics Coding
Free while signed in. Answers cite the passages they came from.

A controlled study pits ChatGPT variants against University of Durham physics students on Python coding assignments, finding that humans still outperform even the strongest prompt-engineered GPT-4.
Setup: 50 student submissions and matched GPT-3.5 / GPT-4 submissions were blind-scored by three independent markers, with prompt engineering applied as an explicit variable.
Students win on average: Students averaged 91.9%, while GPT-4 with prompt engineering scored 81.1% (SE 0.8) - a statistically significant gap.
Prompt engineering matters: Prompt engineering produced large, highly significant improvements for both GPT-3.5 (p ≈ 5×10⁻⁹) and GPT-4 (p ≈ 1.7×10⁻⁴).
Detectable: Human markers correctly identified AI-authored submissions 85.3% of the time, suggesting that while LLM output is close to student quality, it remains stylistically distinguishable.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack