Structure-Aware Calibration Refinement for Monocular Roadside Cameras Using a Horizontal Ground-Coordinate Objective and Road-Patch Phase Correlation
Ikhyeon Jo, Gooman ParkMonocular roadside cameras require accurate image-to-ground mapping and correction of view changes caused by mounting-structure deformation. We combine a horizontal ground-coordinate calibration objective with road-patch phase correlation and affine correction on selected approximately planar road sections with an approximately constant grade. The initial calibration uses 190 reviewed image–global navigation satellite system (GNSS) correspondences from 14 equally weighted cameras at seven sites. With all available control points used for fitting, the camera-balanced mean horizontal position error (MHPE) at those points decreases from 0.259 to 0.196 m (24.4%), and the all-point root mean square error (RMSE) decreases for every camera. This is in-sample fitting performance at the same control points used to estimate the mapping, not independent predictive accuracy. Complementary landmark-level leave-one-out (LOO) evaluation measures prediction at withheld points: the interior LOO decreases from 0.331 to 0.318 m and the exterior LOO from 0.445 to 0.384 m. Overall LOO decreases from 0.398 to 0.348 m, with improvement in seven out of 14 cameras, indicating geometry-dependent predictive benefit. On the original 40 correction frames, Phase–Affine has a lower observed MHPE under the present annotations: 0.392 m, compared with 0.462 m for native-resolution patch DISK + LightGlue with direct robust affine fitting; both yield 40/40 finite evaluations. Their measured mean core times are 9.858 and 106.094 ms, respectively. A 92-frame hourly experiment examines six approved patch centers, five sizes, normalized cross-correlation (NCC) thresholds, patch subsets and transformation models, retaining large finite errors and separately reporting missing transforms. Full-image learned features are more accurate than unfiltered 150 px Phase–Affine on these hourly sequences. The results demonstrate improved in-sample calibration fit and a lower measured core runtime for phase-based correction in the evaluated patch comparison. The small observed MHPE difference is interpreted cautiously because single-annotator repeatability was not formally measured. Held-out and sensitivity analyses delimit these benefits rather than establish uniform superiority across cameras, patch configurations or operating conditions.