Knowledge Fusion of LLMs (FuseLLM)
Free while signed in. Answers cite the passages they came from.

FuseLLM proposes fusing the capabilities of multiple existing LLMs into a single target model by distilling their output distributions rather than retraining from scratch.
Distribution-level fusion: Leverages the generative distributions of multiple source LLMs (e.g., Llama 2, OpenLLaMA, MPT) as fine-grained training signals for the target model via continual training.
Capability transfer: Transfers reasoning, common-sense, and code-generation strengths from heterogeneous sources into a target Llama 2 backbone.
Outperforms individual sources: FuseLLM improves over each source model on downstream benchmarks, suggesting that the fused model captures complementary strengths rather than just averaging.
Cheaper than retraining: Achieves meaningful capability gains with continual training rather than large-scale pretraining, positioning fusion as a practical alternative to training a new LLM.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack