LLMs on Tabular Data: A Survey

A survey that maps how LLMs are being applied to tabular data tasks - a domain historically dominated by gradient-boosted trees and specialized architectures.
Ask this paper
Task coverage: Organizes work across prediction, data synthesis, question answering, and table understanding, giving a single view of a previously scattered literature.
Core techniques: Catalogs prompting strategies (schema serialization, example selection), fine-tuning recipes, and table-specific encoding methods.
Datasets and metrics: Systematizes benchmarks, evaluation metrics, and the dominant models used across tabular-LLM studies.
Open problems: Identifies unexplored directions including robust handling of mixed-type columns, scalable table pretraining, and integration with traditional tabular baselines.