← all papers · overview

Sozkz: Training Efficient Small Language Models For Kazakh From Scratch

Abstract

Kazakh, a Turkic language spoken by over 22 million people, remains underserved by existing multilingual language models, which allocate minimal capacity to low-resource languages and employ tokenizers ill-suited to agglutinative morphology. We present SozKZ, a family of Llama-architecture language models (50M-600M parameters) trained entirely from scratch on 9 billion tokens of Kazakh text with a

Related papers

Ranked by semantic similarity — how closely each paper's abstract matches this one (100% = near-identical topic).