LLM Augmented LLMs (CALM)

Google's CALM composes a large anchor LLM with smaller specialist models via learned cross-attention, unlocking new capabilities without retraining either model.
Ask this paper
Cross-attention composition: Small cross-attention layers learn to route information between the anchor LLM and an augmenting specialist model, combining their representations at multiple depths.
Low-resource language boost: Augmenting PaLM 2-S with a small specialist model trained on low-resource languages materially improves English translation and arithmetic reasoning in those languages.
+40% code gains: Augmenting with a code-specialist model delivers roughly a 40% improvement over the base code model on code generation and explanation tasks.
Modular capability growth: Suggests a future where capability extension happens by composing cheap specialists against an expensive frozen base, rather than retraining the base model on every new domain.