🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 6 – Sep 6, 2026
Data

From Zero to Hero: An Open LLM Ecosystem for Armenian

First page
From Zero to Hero: An Open LLM Ecosystem for Armenian
The curator’s take

Erik Arakelyan and colleagues at NVIDIA and COPA release the first open Armenian LLM published together with the data and recipe needed to reproduce it, and report a data-contamination finding along the way.

Ask this paper

Key points
01

Two released datasets. ArmWeb is 4.37M validated Armenian news documents; ArmSTEM is 373K parallel English-Armenian math and science problems with step-by-step solutions, verified by answer-preserving LLM judgment and by humans.

02

arm-gemma-e4b beats every existing open Armenian model as well as its unadapted Gemma-4-E4B base, after continued pretraining on both datasets.

03

News-only pretraining improves fluency while eroding knowledge, a pattern the authors also find in existing Armenian models, and a small share of verified translated STEM data reverses the loss. That is a directly reusable recipe for other low-resource languages.

04

A contamination result worth noting separately. The largest public Armenian corpora overlap web-derived evaluation panels heavily, including a train/test self-overlap inside FineWeb-2.

05

Everything is released: data, models, and code, which is what distinguishes this from prior Armenian model releases.

Abstract

Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armenian LLM has been released with the data and recipe needed to reproduce it. To address this gap, we curate and release two datasets. ArmWeb is an extensively validated corpus of 4.37M Armenian news documents. ArmSTEM is a parallel English-Armenian collection of 373K math and science problems with step-by-step solutions, translated into Armenian and verified through both answer-preserving LLM judgment and human evaluation. Continued pretraining of Gemma-4-E4B on these datasets yields arm-gemma-e4b, which outperforms every existing open Armenian model as well as its unadapted base, and is the first open Armenian LLM with complete training data and recipe. Our ablations show that news-only continued pretraining improves fluency while eroding knowledge, a pattern we also observe in existing Armenian models, and that a small share of verified translated STEM data reverses the loss. We further find that the largest public Armenian corpora overlap web-derived evaluation panels heavily, including a train/test self-overlap inside FineWeb-2. We openly release all data, models, and code.

Every Monday
Get next week’s papers.
Subscribe on Substack