International Journal of Engineering and Information Systems (IJEAIS)

Title: Automatic Detection Of Spelling Errors In Texts: Methods, Approaches And Algorithms With Application To The Uzbek Language

Authors: Nazira Sobirova

Volume: 10

Issue: 4

Pages: 128-139

Publication Date: 2026/04/28

Abstract:
The rapid expansion of digital communication environments has significantly increased the amount of written textual data produced daily. Ensuring orthographic accuracy in such large-scale text streams has become a critical task in natural language processing (NLP). Automatic spelling error detection systems play an essential role in improving text quality, supporting language learning, enhancing search engine performance, and facilitating intelligent human-computer interaction. While spelling correction technologies are well developed for high-resource languages, morphologically rich and agglutinative languages such as Uzbek still face substantial challenges due to complex word formation processes, extensive suffixation, and phonetic alternations. This study investigates theoretical foundations, computational approaches, and algorithmic solutions for automatic spelling error detection with particular attention to Uzbek-language texts. The paper analyses rule-based, dictionary-driven, statistical, and neural-network-based methods, including edit distance algorithms, probabilistic language models, and transformer architectures. Special emphasis is placed on distinguishing between non-word errors and real-word contextual errors, which remain one of the most difficult problems in automated text processing. A hybrid framework combining linguistic rules, morphological analysis, and contextual neural modelling is proposed. Uzbek-language examples are incorporated to demonstrate algorithm behaviour under real linguistic conditions. Experimental modelling shows that integrating classical algorithms with contextual deep learning significantly improves detection accuracy while reducing false positives. The proposed approach contributes to the development of scalable orthographic analysis tools for low-resource languages and provides methodological guidance for future NLP research in Turkic languages.

Download Full Article (PDF)