🚀NEW LABGetting Started with Claude AgentsStart lab
Training

Knowledge Fusion of LLMs (FuseLLM)

First page
Knowledge Fusion of LLMs (FuseLLM)
Paper summary

FuseLLM proposes fusing the capabilities of multiple existing LLMs into a single target model by distilling their output distributions rather than retraining from scratch.

Ask this paper

Key points
01

Distribution-level fusion: Leverages the generative distributions of multiple source LLMs (e.g., Llama 2, OpenLLaMA, MPT) as fine-grained training signals for the target model via continual training.

02

Capability transfer: Transfers reasoning, common-sense, and code-generation strengths from heterogeneous sources into a target Llama 2 backbone.

03

Outperforms individual sources: FuseLLM improves over each source model on downstream benchmarks, suggesting that the fused model captures complementary strengths rather than just averaging.

04

Cheaper than retraining: Achieves meaningful capability gains with continual training rather than large-scale pretraining, positioning fusion as a practical alternative to training a new LLM.

Every Monday
Get next week’s papers.
Subscribe on Substack