Generative AI, Performance, and Learning: A Framework for Comparing Interventions Across Assessment Regimes
Attila KovariGenerative artificial intelligence (AI) can improve students’ work while assistance is available, but better AI-assisted performance does not necessarily show what students have learned or can do independently. The methodological gap is that outcomes collected under different AI-access, timing, and task conditions are often treated as comparable even when they answer different questions. This paper proposes a framework for making those conditions explicit before intervention effects are compared or synthesized. It separates what is being assessed, the task, the conditions under which the outcome is produced, and evidence used to verify AI use or non-use. A retrospective secondary analysis of Wong and Qiu illustrates the problem. For expert-rated originality, unrestricted ChatGPT-4 use outperformed a learner-first strategy on a stuffed-bunny improvement task (+0.78 points), whereas learner-first outperformed unrestricted use on a vocabulary-game task under the study’s no-AI protocol (−0.73; Holm-adjusted p=0.01351). Usefulness showed the same directional pattern, whereas elaboration did not. These results do not establish general creativity, durable learning, or a causal effect of removing AI because task, sequence, and access conditions also changed. For educators and researchers, the practical implication is that AI-assisted performance, independent performance, retention, and transfer are different outcomes and should not be treated as interchangeable without justification. The proposed Assessment-Regime Reporting Profile provides a compact way to document these conditions before findings are generalized, ranked, or pooled.