🚀NEW LABGetting Started with Claude AgentsStart lab
Memory · Evaluation · Multimodal

Gemini 1.5

First page
Gemini 1.5
Paper summary

Google DeepMind's Gemini 1.5 is a multimodal MoE LLM that scales context to 1M tokens (10M in research settings) while matching or surpassing Gemini 1.0 Ultra on standard benchmarks.

Ask this paper

Key points
01

MoE architecture: A sparsely-activated mixture-of-experts design gives Gemini 1.5 Pro Ultra-class quality at substantially less compute per token.

02

Million-token context: Supports up to 1M tokens in production (10M in research) covering text, video, and audio - enabling reasoning over entire books, hours of video, and multi-hour audio files.

03

Near-perfect retrieval: Achieves >99% accuracy on needle-in-a-haystack retrieval up to at least 10M tokens, significantly beyond contemporary long-context baselines.

04

Quality plus scale: Gemini 1.5 Pro matches or outperforms Gemini 1.0 Ultra on standard benchmarks and sets SoTA on long-document QA, long-video QA, and long-context ASR.

Every Monday
Get next week’s papers.
Subscribe on Substack