DOI: 10.3390/app16168003 ISSN: 2076-3417

Improved Classification-Based Monocular Depth Estimation for Robust Generalization of Diverse Scenes

Wei-Jong Yang, Guo-Wei Wu, Jar-Ferr Yang

Recently, monocular depth estimation methods with neural networks trained using an image-depth database have become widely used and prevalent techniques. However, their generalization abilities are often constrained by training data: models trained on indoor datasets may perform well on other indoor datasets but exhibit some performance degradation for outdoor scenes. Therefore, developing a monocular depth estimation network with strong generalization capability for diverse scenes is challenging. In this paper, we propose a monocular depth estimation model that exhibits high generalizability and performs effectively across various scenarios without additional fine-tuning. In terms of model design, we adopt a powerful visual encoder to extract rich semantic features and incorporate the proposed multiple depth prediction heads with self-adaptability for robust depth estimation. With the designed modules and architectural refinements, the proposed model achieves a favorable balance between parameter efficiency and inference accuracy. After testing several unseen datasets, the simulation results demonstrate that the proposed method achieves better estimation performance than the baseline method and shows better robust visualization performance on multiple seen and unseen datasets.

More from our Archive