Diagnosing Reference-Side Language and Dialect Leakage in Cross-Lingual Tibetan Voice Cloning
Tsehang Dorjee, Duo La, Zhaxi LengbenCross-lingual voice cloning must preserve the reference voice while following the requested language or dialect. We diagnose reference-side interference in dots.tts across Amdo, U-Tsang, and Kham Tibetan; Mandarin; and English. Across the tested conditions, control gains were statistically inconclusive or accompanied by losses in speaker similarity and intelligibility. The primary 147-row display matrix was deduplicated into 120 unique reference–target conditions with five fixed reference prompts. A three-model speaker-disjoint waveform evaluator (SDWE) achieved language-level accuracy of 0.978 and a macro-averaged F1 score (macro-F1) of 0.976; its Tibetan-subset macro-F1 of 0.644 makes dialect findings exploratory. Within this matrix, disabling prompt latent prefill in the baseline system reduced language-level leakage from 0.262 to 0.095, while the crossed-cluster 95% confidence interval, [−0.081,0.450], included zero. Reference x-vector re-pairing showed no stable benefit over a budget-matched continuation; fixed-ratio gating missed the joint control and fidelity criteria; and strong classifier guidance reduced speaker similarity to 0.602 without improving target-dialect hits. A separate prespecified development stress test used 30 tasks, ten distinct reference recordings—two per variety, nine outside the primary matrix—and six target texts. None of nine block-by-time prompt-path masks qualified, and every crossed-cluster confidence interval included zero. An independent Hidden-Unit BERT (HuBERT) layer-9 audit on a source-group-disjoint cohort obtained a dialect macro-F1 of 0.506±0.031 and retained strong source information, so it serves as sensitivity evidence rather than replacing the primary evaluator. A 20-task evaluation by one native Tibetan listener remains exploratory case-level evidence. The combined results identify English-reference-to-Tibetan synthesis as the most concentrated failure route and show that fixed global or block-time interventions do not provide stable target control across the tested prompts and texts.