Challenges and enhancements in Turkish automatic lip reading using deep learning models
Published in Signal, Image and Video Processing, 2026
Recommended citation: Sabaz F., Atıla Ü., Dörterler M., Uçan A., (2026) Challenges and enhancements in Turkish automatic lip reading using deep learning models, Signal, Image and Video Processing, 20(4), 237 https://doi.org/10.1007/s11760-026-05252-2
This study investigates automatic lip reading for Turkish using a deep learning architecture combining CNN and LSTM components. Using 67,080 video samples spanning 111 words and 113 sentences, the study identifies key obstacles in Turkish lip reading through a new “Sentences with Derived Words” dataset that captures the effect of Turkish agglutinative morphology on recognition performance. Bilabial consonant-containing words show superior recognition, while shorter and structurally similar words face higher misidentification risk. The work demonstrates that morphological complexity, acoustic similarity, and word duration substantially affect recognition accuracy, and proposes enhancement strategies based on phonetic feature incorporation and dataset expansion.
Recommended citation: Sabaz F., Atıla Ü., Dörterler M., Uçan A., (2026) Challenges and enhancements in Turkish automatic lip reading using deep learning models, Signal, Image and Video Processing, 20(4), 237