🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal

Chameleon

First page
Chameleon
Paper summary

a family of token-based mixed-modal models for generating images and text in any arbitrary sequence; reports state-of-the-art performance in image captioning and outperforms Llama 2 in text-only tasks and is also competitive with Mixtral 8x7B and Gemini-Pro; exceeds the performance of Gemini Pro and GPT-4V on a new long-form mixed-modal generation evaluation.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack