Transformers as Support Vector Machines
First page

Paper summary
A theoretical paper establishing a formal connection between self-attention optimization and hard-margin SVM problems.
Ask this paper
01
Hard-margin SVM connection: Shows the optimization geometry of self-attention in transformers exhibits a direct connection to hard-margin SVM problems.
02
Implicit regularization: Gradient descent without early stopping leads to implicit regularization, with attention converging toward SVM-like solutions.
03
Theoretical foundation: Provides a rare closed-form theoretical lens on self-attention dynamics, cutting through much of the "transformers as black box" framing.
04
Future analysis tool: The SVM connection gives researchers a principled tool to analyze attention convergence, generalization, and feature selection.