DOI: 10.3390/app16157630 ISSN: 2076-3417

Auditing Per-User Reliability in Webcam-Based Gaze Estimation: Subject- and Pose-Stratified Analysis with an Exploratory Multi-Criteria Decision Scaffold

Kaveti Pavan, Sota Asahara, Nagarajan Ganapathy, Hiroyuki Sugimori

Low-cost webcam gaze estimation is increasingly proposed as a front-end sensor for self-monitoring of attention and visual well-being, yet average accuracy hides where and for whom such a sensor is reliable. Using only three public gaze datasets (MPIIFaceGaze, Columbia Gaze, and RT-GENE; no new data collection), we make three contributions. First, taking a naive sample-wise split (shuffling samples, not subjects) that places every one of 84 subjects across train/validation/test as a deliberate reference baseline, we show apparent error rising from 4.47 deg on seen subjects to 11.63 deg on subject-disjoint unseen subjects (each a five-seed mean; 95% CI: +/−0.11 deg and +/−4.64 deg, respectively), with an optimism of +7.15 deg (~62%); every naive seed lies far below every subject-disjoint seed, and because the unseen estimate varies widely across seeds, its magnitude should be read as indicative rather than precise. Second, we audit reliability per subject and head pose stratum without demographic labels, revealing a 4.9× spread across users and a Monte Carlo dropout uncertainty that provides a partial, borderline per-subject reliability signal (Pearson r = 0.48; Spearman rho = 0.50, with a bootstrap interval including zero); bias-only recalibration lowers error only modestly (subject-macro: 11.79 to 11.50 deg) and does not help the least-reliable users. Third, we convert this measured reliability profile into a per-user decision (trust/recalibrate/abstain) using objective-weight multi-criteria decision making (averaged CRITIC and entropy weights with TOPSIS; the largest weight falls on MC-dropout uncertainty at 0.42, and then pose sensitivity at 0.23); the decision layer disagrees with a single-metric threshold for 35% of users (six of 17) while remaining stable under +/−20% weight perturbation (no action changes). This study provides an exploratory, retrospective, public-data-only implementation of a per-user reliability decision scaffold for low-cost gaze-based self-monitoring.

More from our Archive