How is ChatGPT's Behavior Changing Over Time?
Free while signed in. Answers cite the passages they came from.

Evaluates GPT-3.5 and GPT-4 over months to show significant behavioral drift in deployed systems.
Longitudinal measurement: Compares March vs. June 2023 snapshots of GPT-3.5 and GPT-4 on math, code, sensitive-question answering, and visual reasoning.
Large performance deltas: GPT-4's prime identification accuracy dropped from 97.6% to 2.4% between snapshots, demonstrating drift can be severe and non-monotonic.
Safety and format shifts: Code generation formatting, verbosity, and willingness to answer sensitive questions all changed substantially across versions.
Deployment implications: Highlighted the need for version pinning, regression testing, and behavioral monitoring when building on proprietary APIs - sparking major industry discussion.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack