Zarina Uvalieva, a Machine Learning Engineer at AIRUN Intelligence Lab, has published the scientific paper "Kyrgyz Text Normalization: A Comparative Study of Neural and Rule-Based Approaches", dedicated to advancing artificial intelligence technologies for the Kyrgyz language.
The paper was published as part of the MeLLM Workshop at ACL 2026 (Association for Computational Linguistics), one of the world's most prestigious conferences in natural language processing. ACL holds the highest A* rating in the CORE conference ranking and is widely recognized as one of the leading global venues for research in artificial intelligence and language technologies.
Putting Kyrgyz NLP on the Global Research Map
"ACL is not just another conference. It is one of the top three NLP conferences in the world, alongside EMNLP and NAACL. Papers published there are read by researchers at Google, Meta, OpenAI, Anthropic, and the world's leading universities.
Until now, there have been very few publications from Central Asia in the ACL proceedings, and even fewer from the Kyrgyz Republic. This paper sends an important signal: our region exists on the map of computational linguistics. It opens doors to international collaborations, research grants, invitations to future conferences, internships at top AI labs, and, perhaps most importantly, gives the next generation of Kyrgyz researchers something they can point to and say: 'If they could do it, so can we,'" says Zarina Uvalieva.
The paper is already indexed in Google Scholar, making its findings accessible to the international scientific community.
Advancing Text Normalization for the Kyrgyz Language
The research focuses on one of the fundamental tasks of artificial intelligence: text normalization. In simple terms, it enables AI systems to understand the Kyrgyz language as people actually write it online, with typos, informal spelling, missing punctuation, and inconsistent capitalization.
"What does this mean in practice? If we receive a sentence like 'салам кандайсын кайдасын' without commas, question marks, or capital letters, the system should automatically convert it into 'Салам! Кандайсың? Кайдасың?' before performing any further processing. For English, Russian, and Chinese, such technologies have existed for years. For Kyrgyz, there was nothing," explains Zarina Uvalieva.
"I could have simply taken an existing library and modified it. But that would have been a shortcut with little scientific value. Instead, I chose to go through the entire journey, from fundamental research to a peer-reviewed publication. The Kyrgyz language deserves to be represented in science, not as an afterthought, but as a first-class research subject with its own morphology, challenges, and solutions. Without a scientific foundation, meaningful progress is impossible. Every future study in Kyrgyz NLP can now build upon our dataset and our findings," she adds.
Building the Largest Public Kyrgyz Text Dataset
As part of the research, the AIRUN engineering team created the largest publicly available dataset for the Kyrgyz language to date, containing 1.67 million "noisy-to-clean" text pairs. The researchers also evaluated several state-of-the-art approaches to text normalization. Their fine-tuned model achieved 99.8% accuracy in evaluations conducted by native Kyrgyz speakers and outperformed Google DeepMind's Gemma 4, despite being approximately 32 times smaller.
"Before this publication, any Kyrgyz startup building NLP products faced two choices: either rely on multilingual models that performed poorly on Kyrgyz or spend months collecting their own training data. Now there is a third option: use our open dataset of 1.67 million text pairs together with our open-source model and deploy a production-ready text normalizer within hours. That fundamentally changes the economics of AI development. What previously required a team of five engineers and six months of work can now be accomplished by a single developer in about a week. That means more Kyrgyz AI products, built faster and at a much lower cost," says Zarina.
AIRUN Intelligence Lab Develops Sovereign AI Infrastructure
The research was conducted by AIRUN Intelligence Lab, the research division of BDigital, by a team consisting of Zarina Uvalieva, Bektemir Kumarbai uulu, Adilet Metinov, Tynchtykbek Tashbaltaev, and Nurtilek Alibekov.
Today, AIRUN is the first sovereign AI infrastructure developed in the Kyrgyz Republic. It includes its own Large Language Model (LLM), Automatic Speech Recognition (ASR), Text-to-Speech (TTS), AI Translation, and Digital Avatar technologies. AIRUN solutions are already being deployed across government institutions, the banking sector, media organizations, and private businesses, while also advancing AI technologies for the Kyrgyz language.
Kyrgyz AI Research Gains International Recognition
The publication at ACL 2026 is another demonstration that technologies developed in the Kyrgyz Republic are capable not only of solving practical challenges at home, but also of contributing to the global scientific community in the field of artificial intelligence.
