A Unified Pipeline for Low-Resource Speech Recognition and Understanding: Low-Rank Adaptation, Speaker Diarization, and Graph-Based Retrieval-Augmented Generation
Marija Stojcheva, Goran Petkovski, Igor MishkovskiMacedonian is a low-resource language for automatic speech recognition: annotated speech data are scarce, dialectal variation is substantial, and existing evaluations focus almost entirely on read speech in Standard Macedonian. This paper presents a unified pipeline that converts Macedonian speech, including regional dialects, into accurate transcripts and structured, queryable knowledge, a capability required for applications such as searchable parliamentary archives, broadcast transcription and subtitling, and dialectological documentation. Parameter-efficient adaptation of Whisper large-v3-turbo via low-rank adaptation is evaluated against strong zero-shot and language-specific baselines on four newly curated dialect corpora (Ohrid, Veles, Tikvesh, and Gostivar) and three Standard Macedonian corpora, two of which were collected for this work. The adapted model reduces word error rate by 57–70% relative to the strongest zero-shot baseline and by 33–66% relative to the language-specific BUKI Whisper 2.0 model on dialectal speech, with comparable improvements over zero-shot baselines on standard Macedonian speech, while updating only about 0.7% of parameters. Beyond transcription, the pipeline adds speaker diarization with cross-recording speaker linking and a graph-based retrieval-augmented generation component that enables speaker-, topic-, and time-aware querying of diarized transcripts, evaluated on long-form Macedonian parliamentary recordings. Together, these results establish parameter-efficient adaptation, speaker-aware processing, and graph-based retrieval as a practical and transferable framework for transforming under-resourced speech into accessible, structured knowledge.