Overview of Multilingual LLMs

A first-of-its-kind survey on multilingual LLMs, organized by multilingual alignment principles rather than model-family hierarchy. The authors propose a unified taxonomy and collect open resources to accelerate future research.
Ask this paper
Alignment-first organization: The survey groups methods by how they align languages internally (shared embeddings, cross-lingual pre-training, translation-based alignment, in-context alignment) rather than by model family.
Unified taxonomy: Provides a single framework that covers pretraining-only multilingual models, adapter-based variants, and LLMs adapted post-hoc with translation or code-switching data.
Emerging frontiers: Highlights low-resource languages, cross-lingual transfer for reasoning, and multilingual evaluation as the three frontiers where the community still lacks strong benchmarks.
Open resources: The paper accompanies a curated list of papers, datasets, and leaderboards, lowering the barrier to entry for teams launching new multilingual LLM projects.