🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

Logits of API-Protected LLMs Leak Proprietary Information

First page
Logits of API-Protected LLMs Leak Proprietary Information
Paper summary

The paper shows that the softmax bottleneck in modern LLMs means even logit-level APIs leak enough information to reconstruct hidden architectural details.

Ask this paper

Key points
01

Softmax bottleneck exploit: Because the output distribution is a linear projection of a lower-dimensional embedding, clever API queries can expose the embedding rank and related structural parameters.

02

Cheap recovery: With fewer than $1,000 in API calls, the authors estimate the embedding dimension of OpenAI's gpt-3.5-turbo at ~4,096 and recover related non-public quantities.

03

Auditing use-case: The same technique can detect silent model updates, identify shared ancestry between different provider models, and discover hidden layer widths - making it a potential auditing tool as much as an attack.

04

Mitigations: The paper outlines defenses providers can deploy (logit-bias restrictions, output calibration) while arguing that some transparency gains from this capability are net positive.

Every Monday
Get next week’s papers.
Subscribe on Substack