Patch n' Pack: NaViT
Free while signed in. Answers cite the passages they came from.

A vision transformer handling any aspect ratio and resolution through sequence packing.
Native resolution processing: Packs image patches of arbitrary resolution/aspect-ratio into a single sequence, preserving original information instead of resize-and-crop.
Flexible deployment: Enables compute-quality tradeoffs at inference time without requiring separate models per resolution.
Training efficiency: Sequence packing provides significant training efficiency gains versus fixed-resolution pipelines.
Foundation ViT update: Influenced subsequent multi-modal models (LLaVA, Qwen-VL) that adopted NaViT-style native-resolution image processing.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack