CNN-Based Spatiotemporal Feature Extraction for Video Processing: A Systematic Review
Adrian E. Lopez, Hugo Jimenez-Hernandez, Ana-Marcela Herrera-Navarro, Daniel Canton-Enriquez, Rodrigo Hernandez-Alvarado, Jorge-Luis Perez-Ramos, Arely-Guadalupe Morales-Hernandez, Julio-Cesar Mendez-AvilaThe extraction of spatiotemporal features from video sequences allows for the recognition of actions and the analysis of behaviors in video, making it a key challenge in automated video processing. The literature shows widespread use of deep learning approaches, specifically convolutional neural networks (CNNs). In this context, researchers face the challenge of identifying the advantages, disadvantages, and emerging trends across different architectures, evaluation metrics, and even dataset selection. The objective of this study is to identify the most common CNN architectures, evaluation metrics, datasets, and trends in spatiotemporal feature extraction for video analysis. The selection of articles used the PRISMA methodology and the Joanna Briggs Institute (JBI) methodological framework. From the databases Scopus, Web of Science and the MDPI platform, and based on the inclusion/exclusion criteria, 31 articles that met the criteria were analyzed and synthesized. The search was conducted primarily using the keywords “Convolutional Neural Network,” “video processing,” and “feature extraction,” limiting the selected works to those published between 2020 and the end of 2025. The results mainly show the use of four neural network architectures: 2D CNNs, 3D CNNs, hybrid models (e.g., CNN–RNN, CNN–Transformer, and multi-stream models), and, to a lesser extent, lightweight architectures. Commonly used datasets were identified (e.g., UCF101 and HMDB51). Additionally, standardized evaluation metrics were identified, ranging from accuracy and F1-score to performance measures specific to each case study. The challenges identified center on the heterogeneity of the study datasets, the lack of standardized evaluation metrics, and maintaining a balance between accuracy and computational resource consumption. On the other hand, strong emerging trends toward the use of hybrid models and those integrating transformers have been identified. This systematic review emphasizes the need for clear and robust guidelines that allow for the appropriate selection of CNN architecture, test datasets, and evaluation metrics in applications for extracting spatiotemporal features from video sequences, as well as identifying trends and future lines of research.