Insights on Mitigating Privacy Concerns in Gamification through LLM-Assisted Qualitative Analysis with Minimal Hallucination
Aisvarya Adeseye, Jouni Isoaho, Mohammad TahirGamification improves employee engagement and productivity, but raises concerns about how data is collected, processed, stored, and used. Consequently, analyzing privacy concerns in gamification is complex and time-consuming. Large Language Models (LLMs) are promising because they can understand natural language. However, prompt sensitivity and hallucination affect analytical reliability. Also, data protection risk makes commercial LLMs unsuitable for sensitive contexts. This study examines privacy concerns and mitigation strategies in gamified organizational settings using local LLMs, ensuring in-house data processing while addressing prompt sensitivity and hallucinations. It introduces practical techniques accessible to non-AI experts such as top-k sampling, temperature tuning, prompt optimization, noise reduction, and small-batch processing. The approach was evaluated using interview responses from 82 participants (33 privacy and 49 non-privacy experts) from diverse organizations. Manual analysis with NVivo was compared to LLM-assisted analysis (LLaMA, Gemma, and Phi) for themes, sub-themes, frequency, and impact analysis. Results show that LLM analysis aligns closely with manual results, especially for expert transcripts. However, minor discrepancies appeared in non-expert transcripts due to inconsistent terminology that did not affect the overall strategic outcomes. Ablation studies and larger model tests demonstrate that the proposed strategies generalize and significantly reduce hallucinations. Five main concerns emerged: data collection, usage, security, user awareness, and fairness/trust. Experts prioritized data minimization, anonymization, encryption, and fair algorithms, while non-experts feared surveillance, data misuse, and organizational distrust. The study demonstrates the usefulness of LLMs for early-stage qualitative analysis, proposing a seven-dimensional evaluation framework to assess each model’s performance, strengths, and limitations.