Researchers release an Armenian language model ecosystem

Researchers released COPA, a collection of models, datasets and tools intended to improve Armenian-language AI. The project brings together web text, educational material and a model adapted from Google's Gemma family.

Its ArmWeb corpus contains about 4.37 million documents, equivalent to roughly 3.3 billion tokens under the Gemma tokenizer. ArmSTEM adds approximately 373,000 English–Armenian question pairs, including 324,000 with worked solutions. The arm-gemma-e4b model underwent continued pretraining on 10 billion tokens.

The release includes data-processing code and training and evaluation configurations, making the process easier for other researchers to inspect and reproduce. Usage remains subject to the respective model and dataset licences.

For languages with less training material, better models depend on more than translating a benchmark. This release addresses the data and tooling needed to sustain further work.