🚀NEW LABGetting Started with Claude AgentsStart lab
Training

LLMs on University-Level Physics Coding

First page
LLMs on University-Level Physics Coding
Paper summary

A controlled study pits ChatGPT variants against University of Durham physics students on Python coding assignments, finding that humans still outperform even the strongest prompt-engineered GPT-4.

Ask this paper

Key points
01

Setup: 50 student submissions and matched GPT-3.5 / GPT-4 submissions were blind-scored by three independent markers, with prompt engineering applied as an explicit variable.

02

Students win on average: Students averaged 91.9%, while GPT-4 with prompt engineering scored 81.1% (SE 0.8) - a statistically significant gap.

03

Prompt engineering matters: Prompt engineering produced large, highly significant improvements for both GPT-3.5 (p ≈ 5×10⁻⁹) and GPT-4 (p ≈ 1.7×10⁻⁴).

04

Detectable: Human markers correctly identified AI-authored submissions 85.3% of the time, suggesting that while LLM output is close to student quality, it remains stylistically distinguishable.

Every Monday
Get next week’s papers.
Subscribe on Substack