🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Architecture · Evaluation · Multimodal

Gemini 1.5

Free while signed in. Answers cite the passages they came from.

First page
Gemini 1.5
The curator’s take

Google DeepMind's Gemini 1.5 is a multimodal MoE LLM that scales context to 1M tokens (10M in research settings) while matching or surpassing Gemini 1.0 Ultra on standard benchmarks.

Key points
01

MoE architecture: A sparsely-activated mixture-of-experts design gives Gemini 1.5 Pro Ultra-class quality at substantially less compute per token.

02

Million-token context: Supports up to 1M tokens in production (10M in research) covering text, video, and audio - enabling reasoning over entire books, hours of video, and multi-hour audio files.

03

Near-perfect retrieval: Achieves >99% accuracy on needle-in-a-haystack retrieval up to at least 10M tokens, significantly beyond contemporary long-context baselines.

04

Quality plus scale: Gemini 1.5 Pro matches or outperforms Gemini 1.0 Ultra on standard benchmarks and sets SoTA on long-document QA, long-video QA, and long-context ASR.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack