
Georgian Grammar for Artificial Intelligence is Now Available
Business and Technology University (BTU) has completed the next stage of its Georgian Language Digital Sovereignty project. As part of the initiative, BTU has developed a new guide, “Georgian Grammar for Artificial Intelligence,” designed to support the technological application of the Georgian language’s grammatical system.
The project aims to ensure that, as modern artificial intelligence and language technologies continue to evolve, Georgian is represented not only through textual data but also through a more precise description of its grammatical structure, linguistic relationships, and contextual characteristics.
During the first stage of the project, BTU processed millions of Georgian-language texts, developed digital resources for the Georgian language, and made them openly available on the international platform GitHub. This phase laid the foundation for expanding Georgian-language datasets that can be used by AI systems.
In the second stage, the research shifted toward the structural analysis of the language. Drawing on a Georgian-language corpus of more than 7 billion tokens, researchers examined the language’s grammatical, morphological, and syntactic characteristics. The findings were consolidated into the new guide, which is intended to support more effective processing of Georgian in the development of large language models (LLMs), machine translation systems, and other language technologies.
The work builds on the extensive scientific legacy of Georgian linguists accumulated over many years. BTU’s research represents an effort to interpret and adapt this body of knowledge for contemporary technological applications.
BTU’s artificial intelligence platform, BTUAI, was used in the research process to support the processing and analysis of large volumes of data. The research methodology, evaluation framework, and final interpretation of the findings, however, were defined by Georgian researchers.
According to the authors, the work should not be viewed as a complete computational model of the Georgian language. Rather, it represents an initial step toward the information-based and mathematical modelling of Georgian, creating a foundation for further scientific and technological development.
The first version of the guide is already available through international open-access repositories and can be used by researchers, developers, universities, and organizations working in the field of language technologies.



