SocioTable-KZ: Ethical and Privacy-Aware Bilingual Generation from Sociological Survey Data
Assel Ospan, Madina Mansurova, Zhansaya Zhangabay, Aktoty ZhapparLarge language model (LLM)-based assistants can make large sociological survey collections more accessible, but they may also increase the exposure of respondent-level information and socially sensitive content. This study presents SocioTable-KZ, a privacy-aware, risk-reducing pipeline for bilingual natural-language interpretation of relational sociological survey data from Kazakhstan. The source collection contains 2,385,890 records distributed across 28 linked tables. The pipeline combines identifier exclusion, deterministic JoinGraph serialisation, language-specific QLoRA adaptation of Qwen-family models, a three-class Safe–Sensitive–Unsafe safety module, constrained rewriting, and controlled release. Reference analytical texts were prepared through expert-curated seed examples followed by few-shot candidate generation and factual review. Generation experiments used independent Kazakh and Russian 80/10/10 splits with seed 42. The best reported Qwen3-4B configuration achieved BLEU/ROUGE-L/chrF scores of 29.07/49.88/58.31 for Kazakh and 46.67/65.00/68.40 for Russian. The safety-labelled corpus contained 1347 Kazakh and 3376 Russian instances. On the reported Kazakh evaluation corpus, the safety module achieved 0.925 accuracy and F1-scores of 0.87, 0.95, and 0.70 for Safe, Sensitive, and Unsafe, respectively; Unsafe recall was 0.59. The revised deployment protocol therefore prohibits unattended respondent-level release and requires deterministic identifier screening and human review for non-Safe or uncertain outputs. An official letter from the original data-collecting institution confirms active informed consent and authorised scientific use of the anonymised dataset. The evidence supports preliminary feasibility, not a formal privacy guarantee or production-readiness claim.