ImageBind-LLM
Free while signed in. Answers cite the passages they came from.

Shanghai AI Lab's ImageBind-LLM brings six-modality understanding to LLMs via the ImageBind joint embedding space.
ImageBind backbone: Leverages ImageBind's joint embedding space (covering image, text, audio, depth, thermal, IMU) as a universal multimodal encoder.
Learnable bind network: Aligns ImageBind's visual encoder with a frozen LLM through a learnable bind network, enabling instruction tuning across modalities.
Six-modality input: Responds to instructions over audio, 3D point clouds, video, and beyond - not just text and image.
Generation quality: Maintains high language-generation quality despite the modality diversity, validating the ImageBind-as-bridge approach.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack