🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety

Is Cosine-Similarity Really About Similarity?

Free while signed in. Answers cite the passages they came from.

First page
Is Cosine-Similarity Really About Similarity?
The curator’s take

This paper argues that cosine similarity between learned embeddings does not always measure semantic similarity, and gives analytical examples where it produces arbitrary or non-unique values.

Key points
01

Analytical setting: Studies embeddings derived from regularized linear models, where closed-form expressions expose how cosine similarity depends on regularization choices.

02

Arbitrary "similarities": Different but equally valid optimal solutions can yield wildly different cosine similarities between the same input pairs, undermining the metric's interpretability.

03

Implicit regularization effects: Common deep-learning regularizers induce unintended effects on cosine similarity that practitioners often don't notice or control for.

04

Recommendations: The authors urge caution in using cosine similarity as a universal semantic metric and suggest alternatives such as task-calibrated similarity functions or metric learning.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack