🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

How is ChatGPT's Behavior Changing Over Time?

First page
How is ChatGPT's Behavior Changing Over Time?
Paper summary

Evaluates GPT-3.5 and GPT-4 over months to show significant behavioral drift in deployed systems.

Ask this paper

Key points
01

Longitudinal measurement: Compares March vs. June 2023 snapshots of GPT-3.5 and GPT-4 on math, code, sensitive-question answering, and visual reasoning.

02

Large performance deltas: GPT-4's prime identification accuracy dropped from 97.6% to 2.4% between snapshots, demonstrating drift can be severe and non-monotonic.

03

Safety and format shifts: Code generation formatting, verbosity, and willingness to answer sensitive questions all changed substantially across versions.

04

Deployment implications: Highlighted the need for version pinning, regression testing, and behavioral monitoring when building on proprietary APIs - sparking major industry discussion.

Every Monday
Get next week’s papers.
Subscribe on Substack