DOI: 10.1155/jece/1323497 ISSN: 2090-0147

Adaptive Prompt Engineering for Real‐Time Deepfake Detection in Video Streams Using Multibranch Visual Representation Learning

Mahdi Ajdani

The rapid proliferation of deepfake technologies poses significant threats to media authenticity, privacy, and public trust. While existing deepfake detection models achieve high accuracy in offline scenarios, real‐time detection in continuous video streams remains a major challenge due to latency, a lack of contextual awareness, and adaptive forgery patterns. In this paper, we propose a novel framework that leverages adaptive prompt engineering within a multibranch visual representation learning architecture to detect deepfakes in real‐time video streams. Our method dynamically adjusts visual prompts based on spatiotemporal features, microexpressions, and frame‐wise inconsistencies, guiding the backbone model toward salient forgery cues. The architecture comprises two primary branches: one dedicated to spatial appearance modeling and the other to temporal behavioral patterns. Prompt tokens are optimized during training and refined during inference to respond adaptively to variations in the video stream. We evaluate the proposed framework on benchmark datasets including FaceForensics++, Celeb‐DF, and DFDC, achieving state‐of‐the‐art accuracy (97.3%) with significantly reduced inference latency (< 85 ms per frame) compared to existing methods. Furthermore, ablation studies demonstrate the effectiveness of adaptive prompts in improving both robustness and interpretability in adversarial and unseen scenarios. This study highlights the potential of prompt‐based architectures for practical and scalable deepfake mitigation, paving the way for reliable deployment in real‐world streaming environments such as video conferencing, media platforms, and surveillance systems.