DOI: 10.28945/5824 ISSN: 1547-9714

Bridging Proficiency Gaps: ChatGPT-Enhanced STAD for Achievement, Attitudes, and Equity in Multilingual Higher Education

Yu-Heng Chen, Kimyung Keng

Aim/Purpose: This study examined whether embedding ChatGPT (GPT-4o) within Student Teams Achievement Divisions (STAD) cooperative learning improves academic achievement, learning attitudes, and interaction equity for linguistically heterogeneous undergraduates. Background: Linguistically diverse classrooms face an equity problem in which language proficiency operates as a status characteristic. Most existing AI-in-education studies report aggregate effects rather than subgroup outcomes, and few embed AI within structured cooperative architectures. Methodology: An 18-week quasi-experimental design with three time points was conducted at a Taiwanese university (N = 125; experimental n = 65, control n = 60). Participants were stratified by Test of Chinese as a Foreign Language (TOCFL) proficiency band, with Band A2 as the lower-proficiency threshold. The experimental condition followed a six-phase STAD workflow with tiered prompting and progressive scaffolding fading. Analyses included repeated-measures ANOVA, ANCOVA, and content analysis of reflection logs. Contribution: This study contributes (a) initial subgroup-level evidence, awaiting multi-site replication, that AI-enhanced STAD may narrow proficiency-based achievement gaps in content-area cooperative learning, (b) a replicable governance-embedded design, and (c) a theoretical extension positioning AI as a more knowledgeable other within the zone of proximal development. Findings: Five findings emerged. The experimental group exhibited steeper achievement growth (posttest d = 0.91). Four of six attitude dimensions improved significantly, led by Learning Confidence (d = 0.97). Self-reported prompting purposes differed substantially by proficiency subgroup (Cramér’s V = 0.37), with lower-proficiency learners reporting more language-clarification queries. The proficiency-based achievement gap narrowed from 16.60 to 4.00 points (Group × Proficiency partial η² = 0.050). Exploratory and preliminary evidence further suggested that female students reported higher residual subject anxiety than male students. However, this finding did not survive Bonferroni correction and is constrained by low subscale reliability. Recommendations for Practitioners: Practitioners should embed AI access at specific interaction nodes, differentiate scaffolding through tiered prompts, govern AI use through written agreements and reflection logs, plan progressive fading, and monitor affective outcomes for vulnerable subgroups. Impact on Society: The equity evidence suggests that blanket AI prohibitions may inadvertently disadvantage linguistically diverse learners. The governance-embedded design offers an adaptable model for institutional and national AI-use policies. Future Research: Future research should deploy automated, timestamped AI usage logs; replicate across institutions and disciplines; revise low-reliability subscales using rigorous CFA on larger samples (N ≥ 300); and conduct longitudinal follow-up to assess persistence of equity effects.

More from our Archive