Prompt Sensitivity Under Semantic Perturbations in CLIP-Family Models for Zero-Shot Classroom Behavior Analysis
Yan Ma, Lizhuo Zhang, Xinjie WuVision–language foundation models such as CLIP are increasingly used for zero-shot behavior recognition, yet their robustness to prompt variations remains poorly understood. This paper investigates prompt sensitivity as a critical robustness concern in CLIP-family models for zero-shot classroom behavior analysis, treating prompt wording as a controlled semantic perturbation. Five representative vision–language models (CLIP/OpenAI, OpenCLIP/LAION, SigLIP2, EVA02-CLIP, and DFN-CLIP) are evaluated on three public classroom behavior benchmarks under a strict symmetric protocol. We compare four generic prompt strategies with a training-free Class-Aware Prompt Ensemble (CAPE). Results show that minor prompt changes can cause catastrophic performance degradation. On SigLIP2, an alternative wording of CAPE reduces Hit@1 on TeacherBehavior from 85.5% to 31.4%, a 54.1 percentage-point drop that exceeds the differences between model backbones. Across all five models, action-oriented prompts improve Hit@1 by up to 54 percentage points compared with label-only prompts. We further demonstrate that the apparent superiority of zero-shot CLIP over supervised linear probes largely arises from metric asymmetry. While zero-shot methods achieve higher Hit@1, they consistently underperform linear probes in multi-label evaluation (Sample-F1: 60–66% vs. 88–90%; Macro-F1: 49–60% vs. 68–78%). Bootstrap confidence intervals and paired-bootstrap significance tests further show that several reported performance differences are not statistically significant. These findings reveal prompt sensitivity as a fundamental deployment risk for vision–language foundation models in domain-specific behavior analysis. Prompt variations involving only a few words can silently undermine recognition performance while remaining hidden by conventional evaluation metrics. We therefore recommend that future benchmark studies report prompt configurations, multi-label F1 scores, and uncertainty estimates alongside headline Hit@1 to provide a more complete and reliable assessment of model capability.