🚀NEW LABGetting Started with Claude AgentsStart lab
Data

MAmmoTH2

First page
MAmmoTH2
Paper summary

harvest 10 million naturally existing instruction data from the pre-training web corpus to enhance LLM reasoning; the approach first recalls relevant documents, extracts instruction-response pairs, and then refines the extracted pairs using open-source LLMs; MAmmoTH2-7B's (Mistral) performance increases from 11% to 34% on MATH and from 36% to 67% on GSM8K.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack