Stealing Part of a Production Language Model

The paper demonstrates the first practical attack that extracts the embedding-projection layer of production LLMs through their ordinary logit APIs.
Ask this paper
Attack setup: By querying the API with carefully chosen prompts and analyzing the logit outputs, the attacker can reconstruct the final projection matrix from public API access alone.
Concrete extractions: The attack recovers Ada (hidden dim 1024) and Babbage (hidden dim 2048) matrices for under $20, and estimates GPT-3.5-turbo's hidden dimension for under $2,000.
Cross-provider: Similar attacks apply to other production LLMs including PaLM-2, indicating the vulnerability is structural rather than tied to any one provider.
Mitigations: The authors propose defenses such as restricting logit-bias API features, adding noise, and careful API design to close off the attack surface.