Enhancing Contrastive PU Learning for Ad Fraud Detection with Multi-Granularity Diffusion
Lifei Wei, Weifan Yang, Yan Meng, Xinyu Meng, Le YuDigital ad fraud continuously evolves through device spoofing, behavior simulation, and coordinated traffic attacks, making the detection of sophisticated fraudulent activities a persistent challenge. In real-world settings, only a small fraction of fraudulent traffic receives reliable labels, while vast amounts of unlabeled data may still contain latent fraud. This severely limits the generalization ability of conventional supervised learning models. To address this issue, we propose DiffVC—a novel contrastive learning framework tailored for ad fraud detection under the Positive-Unlabeled (PU) learning paradigm. DiffVC employs a multi-granularity diffusion augmentation strategy that builds upon a denoising diffusion probabilistic model to generate semantically augmented samples at three granularities: weak, medium, and strong. This strategy expands the latent fraud feature space while preserving diverse semantic information. We further introduce a diffusion-distance-calibrated similarity that dynamically adjusts constraints between augmented samples based on their diffusion distance, thereby improving the discrimination of unlabeled fraud. In addition, we design a dual-gated residual Transformer encoder that adaptively captures high-order feature interactions via gated residual connections and a channel recalibration mechanism. Experimental results on multiple real-world ad fraud detection datasets demonstrate that DiffVC consistently outperforms state-of-the-art methods and achieves stronger generalization. Our results confirm the effectiveness and practical applicability of DiffVC for ad fraud detection under label-scarce conditions.