🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture · Multimodal

Meta-Transformer

First page
Meta-Transformer
Paper summary

A unified framework performing learning across 12 different modalities with a shared backbone.

Ask this paper

Key points
01

12-modality coverage: Handles text, image, point cloud, audio, video, X-Ray, infrared, hyperspectral, IMU, graph, tabular, and time-series data.

02

Frozen encoder design: Uses a frozen modality-agnostic encoder paired with modality-specific tokenizers and lightweight task heads.

03

Extreme generality: Demonstrates that a single backbone can serve both fundamental perception and practical applications like medical imaging and industrial sensing.

04

Universal encoder direction: Points toward future architectures where a single foundation model serves as the universal encoder for any modality.

Every Monday
Get next week’s papers.
Subscribe on Substack