Meta-Transformer
Free while signed in. Answers cite the passages they came from.

A unified framework performing learning across 12 different modalities with a shared backbone.
12-modality coverage: Handles text, image, point cloud, audio, video, X-Ray, infrared, hyperspectral, IMU, graph, tabular, and time-series data.
Frozen encoder design: Uses a frozen modality-agnostic encoder paired with modality-specific tokenizers and lightweight task heads.
Extreme generality: Demonstrates that a single backbone can serve both fundamental perception and practical applications like medical imaging and industrial sensing.
Universal encoder direction: Points toward future architectures where a single foundation model serves as the universal encoder for any modality.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack