🚀NEW LABGetting Started with Claude AgentsStart lab
Data

Dynamic Layer Routing in LLMs

First page
Dynamic Layer Routing in LLMs
Paper summary

A retrofittable way to add per-layer routers to frozen LLMs that decide to skip, execute, or repeat each block. Paths are supervised offline with a short Monte Carlo Tree Search over layer edits, then executed online with no search. Improves accuracy on logic and math while saving layers on average.

Ask this paper

Key points
01

The diagram on page 3 shows the per-layer router, its pooling over windows, and how decisions gate the next block.

02

Out-of-domain generalization is strong. Across MMLU, GSM8k, AIME24, TruthfulQA, SQuADv2, GPQA, AGIEval, and PIQA, the average accuracy drop is about 0.85 percentage points while retaining savings.

03

Compared to LayerSkip, ShortGPT, MindSkip, and FlexiDepth, Dr.LLM attains higher average accuracy with far less training data and no base-model changes.

Every Monday
Get next week’s papers.
Subscribe on Substack