🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal · Evaluation

Mini-Gemini

First page
Mini-Gemini
Paper summary

Mini-Gemini enhances vision-language models by adding a second high-resolution visual encoder that refines details without increasing the number of visual tokens consumed by the LLM.

Ask this paper

Key points
01

Dual-encoder design: A primary low-res encoder produces the visual tokens fed to the LLM while a high-res encoder provides patch-level features used to refine those tokens via attention.

02

Token count stays constant: Because refinement happens in feature space rather than by adding tokens, inference cost stays comparable to the base VLM even as visual fidelity improves.

03

Works across scales: The framework supports LLM backbones from 2B to 34B parameters and pairs image understanding with VLM-guided image generation in a single stack.

04

Benchmark leadership: Achieves leading zero-shot scores across several vision-language benchmarks, with the authors reporting results competitive with or surpassing GPT-4 and Gemini on comparable tasks.

Every Monday
Get next week’s papers.
Subscribe on Substack