DOI: 10.3390/su18199952 ISSN: 2071-1050

Assessment Instruments for Generative AI Literacy in Higher Education: A Systematic Review of Measurement Approaches, Psychometric Properties, and Implications for SDG 4

Willy Adauto-Medina, César León-Velarde, Silvia Fernández-Flores, José Antonio Arévalo-Tuesta, Irma Aybar-Bellido, Maritza Arones, Beatriz Caycho-Salas, Adrián Quispe-Andía

Assessment of generative artificial intelligence (GenAI) literacy in higher education draws on both GenAI-specific instruments and broader artificial intelligence (AI) literacy measures, but these evidence sources are not interchangeable. This systematic review examined the GenAI-literacy assessment landscape while distinguishing direct GenAI-specific evidence from general AI-literacy evidence. Following PRISMA 2020 and a registered INPLASY protocol, four electronic information sources were searched for peer-reviewed English-language studies published from 2022 through 26 July 2026. Thirty-one studies met the eligibility criteria: 23 (74.2%) assessed general AI literacy, whereas eight (25.8%) used GenAI-specific instruments. Self-report Likert-type scales predominated (n = 28; 90.3%); performance-based and hybrid approaches were uncommon. Psychometric evidence centered on internal consistency and factor structure, while temporal stability, measurement invariance, external validity, and item-level functioning were rarely examined. Potential implications were mapped to digital competencies, quality education, and responsible AI use, but inclusive learning appeared in only 11 studies (35.5%). Conclusions about GenAI-specific assessment are therefore limited to eight studies. The evidence indicates a need for GenAI-specific hybrid instruments, broader validation, and accessible performance tasks. Such assessment can inform education but constitutes a diagnostic contribution to SDG 4 rather than evidence of its achievement.