Beyond Visual Cues: Impact Analysis of Multi-Duration PCAPs for Deepfake Video Detection Using Network Traffic
Atif Asim, Muhammad Umair, Nauman Mazhar, Mamoona Naveed AsgharThe use of deepfake media, particularly face-swap and face-shifter videos, has proliferated rapidly on social media and online news outlets. Though there are legitimate uses for this synthetic media in certain applications, such as media reporting, there is always the possibility of misleading people, creating distrust of digital information, and even posing security threats. Most existing studies have focused on deepfake detection using either image or video frames; however, these methods are computationally costly and do not perform well in real time. To address these limitations, this work proposes a pipeline for deepfake video detection using network packet analysis. A novel PCAP dataset is constructed by streaming real and manipulated video content over WebRTC and TCP (Port 8080) protocols, and machine learning models are then trained on 48 extracted network-level features to distinguish deepfake traffic from authentic streams. For implementation, four deepfake video datasets of varying quality are utilized, namely, HIDF, FaceForensics++, SDFVD-V2, and ManualFake-2022, each containing both real and manipulated samples. Videos are segmented into 3, 6, and 9-s clips and sequentially streamed over WebRTC and TCP (Port 8080) protocols to capture network traffic in PCAP format. Experiments are conducted using six classical machine learning classifiers, namely KNN, Logistic Regression, Decision Tree, Random Forest, Histogram Gradient Boosting, and Naïve Bayes, trained on 48 network-level features. The proposed pipeline achieved the highest classification accuracy of 91.2% on TCP (Port 8080) and 79.9% on the WebRTC protocol, both obtained using 6-s video captures.