🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Architecture · Multimodal

Patch n' Pack: NaViT

Free while signed in. Answers cite the passages they came from.

First page
Patch n' Pack: NaViT
The curator’s take

A vision transformer handling any aspect ratio and resolution through sequence packing.

Key points
01

Native resolution processing: Packs image patches of arbitrary resolution/aspect-ratio into a single sequence, preserving original information instead of resize-and-crop.

02

Flexible deployment: Enables compute-quality tradeoffs at inference time without requiring separate models per resolution.

03

Training efficiency: Sequence packing provides significant training efficiency gains versus fixed-resolution pipelines.

04

Foundation ViT update: Influenced subsequent multi-modal models (LLaVA, Qwen-VL) that adopted NaViT-style native-resolution image processing.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack