🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture · Multimodal

Patch n' Pack: NaViT

First page
Patch n' Pack: NaViT
Paper summary

A vision transformer handling any aspect ratio and resolution through sequence packing.

Ask this paper

Key points
01

Native resolution processing: Packs image patches of arbitrary resolution/aspect-ratio into a single sequence, preserving original information instead of resize-and-crop.

02

Flexible deployment: Enables compute-quality tradeoffs at inference time without requiring separate models per resolution.

03

Training efficiency: Sequence packing provides significant training efficiency gains versus fixed-resolution pipelines.

04

Foundation ViT update: Influenced subsequent multi-modal models (LLaVA, Qwen-VL) that adopted NaViT-style native-resolution image processing.

Every Monday
Get next week’s papers.
Subscribe on Substack