Adapting Large-Scale Foundation Models for Turkic Speech-to-Speech Translation: Fine-Tuned Cascade and Direct Approaches
Aidana Karibayeva, Vladislav Karyukin, Oleg Myssov, Dina Amirova, Balzhan Abduali, Adina KarybayevaThis paper presents a novel approach to speech-to-speech (STS) translation for low-resource Turkic languages. Today, STS has progressed rapidly for high-resource languages; the Turkic family remains significantly underrepresented. Consequently, it is quite challenging to develop a reliable speech translation system for these languages. To address this issue, we have developed two speech translation systems (STS) specifically tailored to Turkic languages with limited resources. The first, TurkicCascadeSTS, is a cascaded system that combines a fine-tuned Whisper-medium speech recognition model, GPT translation, and separate speech synthesis models for each language. The second system is a direct speech translation model based on a fine-tuned SeamlessM4Tv2. Both systems have been tested for translation into Turkic languages, using 24,656 audio recordings per language. The TurkicCascadeSTS system delivered far better results: the average BLEU score rose from 4.68 to 30.60; the METEOR score rose from 15.03 to 44.42; and the word error rate (WER) also fell significantly. These improvements are due to the fact that each module of the system was individually fine-tuned to account for the specific characteristics of each language. Although SeamlessM4Tv2 sometimes produces clearer audio, TurkicCascadeSTS generally delivers higher speech and translation quality for all language pairs. This demonstrates that modular, specially tuned systems are an effective solution for translation into Turkic languages, particularly given their complex structure and limited linguistic resources. Such systems could benefit more than 200 million native speakers of Turkic languages.