Statistical Learning Theory for Inverse-Probability-Weighted Conditional U-Statistics via Delta Sequences Under Functional Missing-at-Random Models
Salim BouzebdaThis paper develops a unified asymptotic theory for inverse-probability-weighted conditional U-statistics of arbitrary fixed order in the presence of missing-at-random responses and infinite-dimensional functional covariates. The target is a conditional higher-order functional generated by a measurable response kernel and evaluated locally on a separable Banach space. Localization is formulated through delta sequences, providing a common framework for kernel, partition, regressogram, orthogonal series, and related smoothing procedures without recourse to finite-dimensional density arguments. For bounded kernels, we establish uniform almost-complete convergence over pseudo-compact functional domains and obtain a sharp decomposition into deterministic localization bias and stochastic fluctuation. The latter is governed by the localized-kernel variance, the envelope of the delta sequence, the metric complexity of the indexing domain, and the small-ball concentration of the functional covariate. Unbounded kernels are treated under explicit weighted moment, truncation, and summability conditions. The feasible theory quantifies the additional perturbation induced by estimating the propensity score and identifies conditions under which this first-stage uncertainty is asymptotically negligible. Pointwise distributional theory is derived through a denominator linearization combined with the Hoeffding decomposition of the centered localized kernel. The Gaussian limit is driven by the first projection, while the higher-order canonical components are shown to be negligible under explicit local-mass, moment, and noncancellation assumptions. This yields oracle-equivalent feasible inference, a consistent first-projection variance estimator, and asymptotically valid studentized confidence intervals. A finite-grid adaptive comparison principle is also developed for data-driven resolution selection. The scope of the theory is illustrated through conditional rank functionals, discrimination with incomplete labels, metric-learning criteria, and functional prediction. Synthetic and semi-synthetic studies based on functional classification, phoneme log-periodograms, and growth trajectories document the finite-sample interaction between covariate-dependent label observation, local information loss, propensity estimation, and inverse-weighting variance.