Knowledge Fusion of LLMs (FuseLLM)

FuseLLM proposes fusing the capabilities of multiple existing LLMs into a single target model by distilling their output distributions rather than retraining from scratch.
Ask this paper
Distribution-level fusion: Leverages the generative distributions of multiple source LLMs (e.g., Llama 2, OpenLLaMA, MPT) as fine-grained training signals for the target model via continual training.
Capability transfer: Transfers reasoning, common-sense, and code-generation strengths from heterogeneous sources into a target Llama 2 backbone.
Outperforms individual sources: FuseLLM improves over each source model on downstream benchmarks, suggesting that the fused model captures complementary strengths rather than just averaging.
Cheaper than retraining: Achieves meaningful capability gains with continual training rather than large-scale pretraining, positioning fusion as a practical alternative to training a new LLM.