ImageBind-LLM
First page

Paper summary
Shanghai AI Lab's ImageBind-LLM brings six-modality understanding to LLMs via the ImageBind joint embedding space.
Ask this paper
01
ImageBind backbone: Leverages ImageBind's joint embedding space (covering image, text, audio, depth, thermal, IMU) as a universal multimodal encoder.
02
Learnable bind network: Aligns ImageBind's visual encoder with a frozen LLM through a learnable bind network, enabling instruction tuning across modalities.
03
Six-modality input: Responds to instructions over audio, 3D point clouds, video, and beyond - not just text and image.
04
Generation quality: Maintains high language-generation quality despite the modality diversity, validating the ImageBind-as-bridge approach.