LLMs on University-Level Physics Coding

A controlled study pits ChatGPT variants against University of Durham physics students on Python coding assignments, finding that humans still outperform even the strongest prompt-engineered GPT-4.
Ask this paper
Setup: 50 student submissions and matched GPT-3.5 / GPT-4 submissions were blind-scored by three independent markers, with prompt engineering applied as an explicit variable.
Students win on average: Students averaged 91.9%, while GPT-4 with prompt engineering scored 81.1% (SE 0.8) - a statistically significant gap.
Prompt engineering matters: Prompt engineering produced large, highly significant improvements for both GPT-3.5 (p ≈ 5×10⁻⁹) and GPT-4 (p ≈ 1.7×10⁻⁴).
Detectable: Human markers correctly identified AI-authored submissions 85.3% of the time, suggesting that while LLM output is close to student quality, it remains stylistically distinguishable.