DOI: 10.1111/1556-4029.70437 ISSN: 0022-1198

Explainable copy‐move forgery detection in videos: Generative adversarial network‐based transfer learning and multi‐aspect transformer with shuffle attention

M. Raghavendra Reddy, C. Anbu Ananth, B. Santhosh Kumar

Abstract

With the nascent reliance on unmanned aerial systems as a means of surveillance, security, and monitoring, it is essential to ensure that video data legitimacy is assured. Object‐based Copy‐Move Forgery (CMF), in which an object is either copied or moved in the frame or across a frame to alter the visual description, is one of the major threats to video integrity. Such manipulations are a problem for traditional detection methods because object motion, occlusion, and irregular background are complex. This research presents an interpretable and robust deep learning architecture to identify such forgeries, that is, a combination of transfer learning with a new Multi‐Aspect Fossa Graph Transformer enhanced with Shuffle Attention (MAFGTN‐SA). To achieve this, a better generative adversarial network, which is built upon GoogLeNet, is used to extract deep spatial‐semantic features using normalized grayscale video frames. These characteristics are then represented by MAFGTN‐SA to represent complex object relationships between the frames and spatial inconsistencies to accurately classify genuine and tampered frames. In order to increase interpretability, the model incorporates Local Interpretable Model‐Agnostic Explanations (LIME) for saliency‐based visualization of tampered regions. Experimental evaluations on three benchmark datasets, SULFA, GRIP, VTD, CG‐1050 v2.0, and COVERAGE, validate the effectiveness of the proposed method, attaining an accuracy of 98.73%, 96.78%, 98.45%, 97.92%, and 97.36%, respectively, and with high F1‐scores of 98.73%, 98.5%, 97.98%, 97.41%, and 96.93%. The results highlight the framework's superior detection capability, reliability, and explainability across diverse video forgery scenarios.

More from our Archive