A Data Mining Framework for Pesticide Co-Occurrence Analysis Using Association Rules and Hamming Distance in Surface Water Monitoring
Tereza Motúzová, Vojtěch UherMonitoring pesticide contamination in surface waters produces datasets with complex co-occurrence patterns that conventional statistical methods may struggle to interpret. This paper presents a data-mining framework for identifying pesticide relationships from binary occurrence data collected through long-term environmental monitoring. Such data can be collected easily and inexpensively and are well suited to a moderate number of samples. The workflow combines data preprocessing, Hamming-distance similarity analysis, frequent itemset mining, and association rule mining using support, confidence, lift, and conviction metrics. It was evaluated using an eighteen-month campaign across seven surface-water localities in the Czech Republic, monitoring 95 pesticides and metabolites, of which 65 were detected. Similarity analysis distinguished pesticide occurrence profiles and identified one locality with a unique contamination composition. Association rule mining generated 9071 pairwise relationships, revealing frequent and locality-specific pesticide combinations. The results also show that high detection density strongly affects conventional association-rule metrics, highlighting the need to interpret confidence and lift alongside occurrence prevalence. The framework offers an interpretable, computationally efficient method for exploring environmental occurrence data of moderate dimension and generating hypotheses about pollution sources, agricultural practices, and transport pathways. It is readily transferable to other monitoring programs based on binary detection records.