LLM360

LLM360 is a framework for fully transparent open-source LLM development, with everything from data to training dynamics released.
Ask this paper
End-to-end transparency: Ships training code, the pretraining corpus, intermediate checkpoints, evaluation code, and analyses - going well beyond the "just weights" openness of earlier "open" LLMs.
Two 7B models: Releases AMBER (general) and CRYSTALCODER (code-specialized) 7B models pretrained from scratch under the framework.
Enables training-dynamics research: Intermediate checkpoints let researchers study loss trajectories, emergent capabilities, and data-effect ablations - typically only possible inside frontier labs.
Standard for openness: Pushes the community's definition of "open-source LLM" from weights to a full training-pipeline standard.