🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal · Retrieval · Evaluation

UniIR

First page
UniIR
Paper summary

UniIR is a unified instruction-guided multimodal retriever that handles eight retrieval tasks across modalities with a single model.

Ask this paper

Key points
01

Instruction-guided: A single retriever conditioned on natural-language instructions determines which retrieval task to perform, rather than one retriever per task.

02

Eight tasks: Handles image-to-text, text-to-image, composed-image retrieval, video retrieval, and other multimodal variants under one umbrella.

03

Zero-shot generalization: Generalizes to unseen retrieval tasks not explicitly trained on, approaching a truly general multimodal retrieval model.

04

M-BEIR benchmark: Ships with a new multimodal retrieval benchmark (M-BEIR) designed to standardize evaluation across tasks and modalities.

Every Monday
Get next week’s papers.
Subscribe on Substack