Big data life cycle management practices, experiences and challenges in computer vision: implications for research data services
Zonghan Lei, Wei Zakharov, Siqing WeiPurpose
Advances in computer vision heavily depend on large data sets, making ethical research data management (RDM) increasingly crucial. This paper, an interdisciplinary study, aims to examine RDM practices and challenges among computer vision scholars and explores a GenAI-assisted qualitative analysis workflow as a methodological contribution.
Design/methodology/approach
Through qualitative case studies of six scholars from five research intensive universities, the authors investigate their data management strategies and challenges. The primary analytical approach involves manual thematic coding with NVivo, complemented by ChatGPT 4o (Qualibot) as an AI-assisted verification tool and a comparative triangulation of both methods to ensure consistency and trustworthiness.
Findings
Key RDM themes identified include data types and sizes, storage and management, sharing and accessibility, services and training and data management struggles. Drawing on the experiences of six computer vision scholars, the findings reveal a strong reliance on hybrid storage systems combining cloud and local storage for large-scale visual data sets. These experience-based insights suggest that major challenges include financial and logistical burdens of data storage, backup, sharing and the need for data management training.
Practical implications
The findings exemplify the need for better institutional support, scalable tools and standardized protocols, offering crucial insights for library and information science professionals developing research data services tailored to big data environments.
Originality/value
This interdisciplinary study provides one of the few empirical, qualitative account of RDM practices specifically within the computer vision community. Methodologically, the authors introduce a novel GenAI-assisted qualitative analysis workflow (Qualibot) rigorously evaluated against manual coding as a complementary validation tool, providing a replicable model for AI-assisted qualitative research in data-intensive domains.