Self-supervised monocular depth estimation in dynamic scenes based on deep learning

  • role: First author第一作者
  • Affiliation:

    Information Engineering University, Institute of Geospatial Information, Zhengzhou 450001, China

  • Email:chengbinbin990816@163.com
  • Introduction:E-mail chengbinbin990816@163.com
CHENG Binbin,  
  • role: Corresponding author通信作者
  • Affiliation:

    Information Engineering University, Institute of Geospatial Information, Zhengzhou 450001, China

  • Email:yuying559101@163.com
  • Introduction:E-mail yuying559101@163.com
YU Ying*,  
  • Affiliation:

    Information Engineering University, Institute of Geospatial Information, Zhengzhou 450001, China

ZHANG Lei,  
  • Affiliation:

    Information Engineering University, Institute of Geospatial Information, Zhengzhou 450001, China

WANG Ziquan,  
  • Affiliation:

    Information Engineering University, Institute of Geospatial Information, Zhengzhou 450001, China

JIANG Zhipeng

реферат

In the real world, completely static scenes do not exist. Monocular depth estimation in dynamic scenes refers to obtaining depth information of dynamic foreground and static background from a single image, which has advantages over traditional stereo estimation methods in terms of flexibility and cost-effectiveness. It has strong research relevance and broad development prospects, playing a key role in downstream tasks, such as 3D reconstruction and autonomous driving. With the rapid development of deep learning technology self-supervised learning without using real data labels has attracted the enthusiasm of many scholars. Many local and foreign scholars have proposed a series of self-supervised monocular depth estimation algorithms to deal with dynamic objects in scenes, laying the research foundation for researchers in related fields. However, a comprehensive analysis of the above methods has yet to be conducted. To address this issue, this study systematically reviews and summarizes the progress of self-supervised monocular depth estimation in dynamic scenes based on deep learning.First, the basic models of self-supervised monocular depth estimation based on deep learning are summarized, and how self-supervised constraints are applied between images is analyzed and explained. Moreover, a basic framework diagram of self-supervised monocular depth estimation based on continuous frames is drawn. The effect of dynamic objects on images is explained from four aspects: epipolar lines, triangulation, fundamental matrix estimation, and reprojection error.Second, commonly used datasets and evaluation metrics for monocular depth estimation research are introduced. The KITTI and Cityscapes datasets provide continuous outdoor image data, while the NYU Depth V2 dataset provides indoor dynamic scene data, which are generally used for model training. The Make3D dataset has depth data but discontinuous images, which are generally used to test the generalization ability of the model. The algorithms are quantitatively analyzed using Root Mean Square Error (RMSE), logarithmic root mean square error (RMSE log), absolute relative error (Abs Rel), squared relative error (Sq Rel), and accuracies (Acc), and the performance of classic monocular depth estimation models in dynamic scenes is compared and analyzed.Then, on the basis of different ways of handling dynamic objects, the research directions of robust depth estimation in dynamic scenes and dynamic object tracking and depth estimation are summarized and analyzed. Dynamic objects are extracted and treated as outliers during training model to minimize their effect, training solely on static background information, which is referred to as robust depth estimation in dynamic scenes. Accurately distinguishing dynamic foreground and static background and processing the two regions separately is referred to as dynamic object tracking and depth estimation. Various algorithms for detecting and segmenting dynamic objects based on optical flow information, semantic information, and other information while estimating their motion are explained. At the same time, the advantages and disadvantages of each type of algorithm are summarized and analyzed on the basis of commonly used evaluation criteria.Finally, the future development directions of monocular depth estimation in dynamic scenes are discussed from the aspects of network model optimization, online learning and generalization, real-time operation capability of embedded devices, and domain adaptation of self-supervised learning.

ключеви́че слова́

remote sensing;dynamic scenes;monocular depth estimation;self-supervised learning;deep learning;3D reconstruction

References

  1. 1.
    Bhat S F, Alhashim I and Wonka P. 2021. AdaBins: depth estimation using adaptive bins//Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE: 4008-4017
  2. 2.
    Bian J W, Zhan H Y, Wang N Y, Li Z C, Zhang L, Shen C H, Cheng M M and Reid I. 2021. Unsupervised scale-consistent depth learning from video. International Journal of Computer Vision, 129(9): 2548-2564
  3. 3.
    Bian J W, Zhan H Y, Wang N Y, Chin T J, Shen C H and Reid I. 2022. Auto-rectify network for unsupervised indoor depth estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(12): 9802-9813
  4. 4.
    Boulahbal H E, Voicila A and Comport A I. 2022. Instance-aware multi-object self-supervision for monocular depth prediction. IEEE Robotics and Automation Letters, 7(4): 10962-10968
  5. 5.
    Casser V, Pirk S, Mahjourian R and Angelova A. 2019. Depth prediction without the sensors: leveraging structure for unsupervised learning from monocular videos//Proceedings of the 33rd AAAI Conference on Artificial Intelligence. Honolulu: AAAI: 8001-8008
  6. 6.
    Cordts M, Omran M, Ramos S, Rehfeld T, Enzweiler M, Benenson R, Franke U, Roth S and Schiele B. 2016. The cityscapes dataset for semantic urban scene understanding//Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas: IEEE: 3213-3223
  7. 7.
    Dai Q, Patil V, Hecker S, Dai D X, van Gool L and Schindler K. 2020. Self-supervised object motion and depth estimation from video//Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Seattle: IEEE: 4326-4334
  8. 8.
    Dong X S, Garratt M A, Anavatti S G and Abbass H A. 2022. Towards real-time monocular depth estimation for robotics: a survey. IEEE Transactions on Intelligent Transportation Systems, 23(10): 16940-16961
  9. 9.
    Eigen D and Fergus R. 2015. Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture//2015 IEEE International Conference on Computer Vision (ICCV). Santiago: IEEE: 2650-2658
  10. 10.
    Eigen D, Puhrsch C and Fergus R. 2014. Depth map prediction from a single image using a multi-scale deep network//Proceedings of the 27th International Conference on Neural Information Processing Systems. Montreal: MIT Press: 2366-2374
  11. 11.
    Feng Z Y, Yang L, Jing L L, Wang H Y, Tian Y L and Li B. 2022. Disentangling object motion and occlusion for unsupervised multi-frame monocular depth//17th European Conference on Computer Vision. Tel Aviv: Springer: 228-244
  12. 12.
    Gao F, Yu J C, Shen H, Wang Y and Yang H Z. 2021. Attentional separation-and-aggregation network for self-supervised depth-pose learning in dynamic scenes//Proceedings of the 4th Conference on Robot Learning. Cambridge: PMLR: 2195-2205
  13. 13.
    Gao X B, Shi X H, Ge Q F and Chen Q Y. 2021. A survey of visual SLAM for scenes with dynamic objects. Robot, 43(6): 733-750
  14. 14.
    Garg R, Kumar V, Carneiro G and Reid I. 2016. Unsupervised CNN for single view depth estimation: geometry to the rescue//14th European Conference on Computer Vision. Amsterdam: Springer: 740-756
  15. 15.
    Gasperini S, Koch P, Dallabetta V, Navab N, Busam B and Tombari F. 2021. R4Dyn: exploring radar for self-supervised monocular depth estimation of dynamic scenes//2021 International Conference on 3D Vision (3DV). London: IEEE: 751-760
  16. 16.
    Geiger A, Lenz P and Urtasun R. 2012. Are we ready for autonomous driving? The Kitti vision benchmark suite//2012 IEEE Conference on Computer Vision and Pattern Recognition. Providence: IEEE: 3354-3361
  17. 17.
    Godard C, Mac Aodha O and Brostow G J. 2017. Unsupervised monocular depth estimation with left-right consistency//Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu: IEEE: 6602-6611
  18. 18.
    Godard C, Mac Aodha O, Firman M and Brostow G. 2019. Digging into self-supervised monocular depth estimation//2019 IEEE/CVF International Conference on Computer Vision (ICCV). Seoul: IEEE: 3827-3837
  19. 19.
    Gordon A, Li H H, Jonschkowski R and Angelova A. 2019. Depth from videos in the wild: unsupervised monocular depth learning from unknown cameras//2019 IEEE/CVF International Conference on Computer Vision (ICCV). Seoul: IEEE: 8976-8985
  20. 20.
    Guizilini V, Ambruş R, Pillai S, Raventos A and Gaidon A. 2020a. 3D packing for self-supervised monocular depth estimation//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle: IEEE: 2482-2491
  21. 21.
    Guizilini V, Hou R, Li J, Ambru R and Gaidon A. 2020b. Semantically-guided representation learning for self-supervised monocular depth. arXiv:2002.12319
  22. 22.
    He K M, Gkioxari G, Dollár P and Girshick R. 2020. Mask R-CNN. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(2): 386-397
  23. 23.
    He K M, Zhang X Y, Ren S Q and Sun J. 2016. Deep residual learning for image recognition//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas: IEEE: 770-778
  24. 24.
    Hui T W. 2022. RM-Depth: unsupervised learning of recurrent monocular depth in dynamic scenes//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans: IEEE: 1665-1674
  25. 25.
    Ilg E, Mayer N, Saikia T, Keuper M, Dosovitskiy A and Brox T. 2017. FlowNet 2.0: evolution of optical flow estimation with deep networks//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu: IEEE: 1647-1655
  26. 26.
    Jiang J J, Li Z Y and Liu X M. 2022. Deep learning based monocular depth estimation: a survey. Chinese Journal of Computers, 45(6): 1276-1307
  27. 27.
    Jung D, Choi J, Lee Y, Kim D, Kim C, Manocha D and Lee D. 2021. DnD: dense depth estimation in crowded dynamic indoor scenes//2021 IEEE/CVF International Conference on Computer Vision (ICCV). Montreal: IEEE: 12777-12787
  28. 28.
    Kang J Z,Wang G Z,He G J,Wang H H,Yin R ,Jiang W,Zhang Z M. 2020. Moving vehicle detection for remote sensing satellite video. Journal of Remote Sensing (in Chinese),24(9): 1099-1107
  29. 29.
    Khan F, Salahuddin S and Javidnia H. 2020. Deep learning-based monocular depth estimation methods—A state-of-the-art review. Sensors, 20(8): 2272
  30. 30.
    Klingner M, Termöhlen J A, Mikolajczyk J and Fingscheidt T. 2020. Self-supervised monocular depth estimation: solving the dynamic object problem by semantic guidance//16th European Conference on Computer Vision. Glasgow: Springer: 582-600
  31. 31.
    Lee S, Im S, Lin S and Kweon I S. 2021a. Learning monocular depth in dynamic scenes via instance-aware projection consistency//35th AAAI Conference on Artificial Intelligence. Virtually: AAAI: 1863-1872
  32. 32.
    Lee S, Rameau F, Pan F and Kweon I S. 2021b. Attentive and contrastive learning for joint depth and motion field estimation//2021 IEEE/CVF International Conference on Computer Vision (ICCV). Montreal: IEEE: 4842-4851
  33. 33.
    Li H H, Gordon A, Zhao H, Casser V and Angelova A. 2021. Unsupervised monocular depth learning in dynamic scenes//Proceedings of the 4th Conference on Robot Learning. Cambridge: PMLR: 1908-1917
  34. 34.
    Li Y M,Guo Q H,Wan B,Qin H N,Wang D Z,Xu K X,Song S L,Sun Q H,Zhao X X,Yang M H,Wu X Y,Wei D J,Hu T Y and Su Y J. 2021. Current status and prospect of three-dimensional dynamic monitoring of natural resources based on LiDAR. National Remote Sensing Bulletin, 25(1):381-402
  35. 35.
    Luo C X, Yang Z H, Wang P, Wang Y, Xu W, Nevatia R and Yuille A. 2020. Every pixel counts ++: joint learning of geometry and motion with 3D holistic understanding. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(10): 2624-2641
  36. 36.
    Marr D and Poggio T. 1979. A computational theory of human stereo vision. Proceedings of the Royal Society of London. Series B, Biological Sciences, 204(1156): 301-328
  37. 37.
    Mayer N, Ilg E, Häusser P, Fischer P, Cremers D, Dosovitskiy A and Brox T. 2016. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas: IEEE: 4040-4048
  38. 38.
    Ming Y, Meng X Y, Fan C X and Yu H. 2021. Deep learning for monocular depth estimation: a review. Neurocomputing, 438: 14-33
  39. 39.
    Ranftl R, Vineet V, Chen Q F and Koltun V. 2016. Dense monocular depth estimation in complex dynamic scenes//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas: IEEE: 4058-4066
  40. 40.
    Ranjan A, Jampani V, Balles L, Kim K, Sun D Q, Wulff J and Black M J. 2019. Competitive collaboration: joint unsupervised learning of depth, camera motion, optical flow and motion segmentation//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach: IEEE: 12232-12241
  41. 41.
    Russakovsky O, Deng J, Su H, Krause J, Satheesh S, Ma S A, Huang Z H, Karpathy A, Khosla A, Bernstein M, Berg A C and Fei-Fei L. 2015. ImageNet large scale visual recognition challenge. International Journal of Computer Vision, 115(3): 211-252
  42. 42.
    Ronneberger O, Fischer P and Brox T. 2015. U-net: convolutional networks for biomedical image segmentation//18th International Conference on Medical Image Computing and Computer-Assisted Intervention. Munich: Springer: 234-241
  43. 43.
    Russell C, Yu R and Agapito L. 2014. Video pop-up: monocular 3D reconstruction of dynamic scenes//13th European Conference on Computer Vision. Zurich: Springer: 583-598
  44. 44.
    Saputra M R U, Markham A and Trigoni N. 2018. Visual SLAM and structure from motion in dynamic environments: a survey. ACM Computing Surveys, 51(2): 37
  45. 45.
    Saunders K, Vogiatzis G and Manso L J. 2023. Dyna-DM: dynamic object-aware self-supervised monocular depth maps. arXiv:2206.03799
  46. 46.
    Saxena A, Sun M and Ng A Y. 2009. Make3D: learning 3D scene structure from a single still image. IEEE Transactions on Pattern Analysis and Machine Intelligence, 31(5): 824-840
  47. 47.
    Schönberger J L and Frahm J M. 2016. Structure-from-motion revisited//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas: IEEE: 4104-4113
  48. 48.
    Silberman N, Hoiem D, Kohli P and Fergus R. 2012. Indoor segmentation and support inference from RGBD images//12th European Conference on Computer Vision. Florence: Springer: 746-760
  49. 49.
    Sun L B, Bian J W, Zhan H Y, Yin W, Reid I and Shen C H. 2024. SC-DepthV3: robust self-supervised monocular depth estimation for dynamic scenes. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(1): 497-508
  50. 50.
    Ummenhofer B, Zhou H Z, Uhrig J, Mayer N, Ilg E, Dosovitskiy A and Brox T. 2017. DeMoN: depth and motion network for learning monocular stereo//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu: IEEE: 5622-5631
  51. 51.
    Varma A, Chawla H, Zonooz B and Arani E. 2022. Transformers in self-supervised monocular depth estimation with unknown camera intrinsics. arXiv:2202.03131
  52. 52.
    Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Kaiser L and Polosukhin I. 2023. Attention is all you need. arXiv:1706.03762
  53. 53.
    Vijayanarasimhan S, Ricco S, Schmid C, Sukthankar R and Fragkiadaki K. 2017. SfM-Net: learning of structure and motion from video. arXiv:1704.07804
  54. 54.
    Wang C Y, Buenaposada J M, Zhu R and Lucey S. 2018. Learning depth from monocular videos using direct methods//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE: 2022-2030
  55. 55.
    Wang P, Shen X H, Lin Z, Cohen S, Price B and Yuille A. 2015. Towards unified depth and semantic prediction from a single image//2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Boston: IEEE: 2800-2809
  56. 56.
    Wang R, Pizer S M and Frahm J M. 2019. Recurrent neural network for (un-)supervised learning of monocular video visual odometry and depth//Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Long Beach: IEEE: 5550-5559
  57. 57.
    Wimbauer F, Yang N, von Stumberg L, Zeller N and Cremers D. 2021. MonoRec: semi-supervised dense reconstruction in dynamic environments from a single moving camera//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville: IEEE: 6108-6118
  58. 58.
    Xu H F, Zheng J M, Cai J F and Zhang J Y. 2019. Region deformer networks for unsupervised depth estimation from unconstrained monocular videos. arXiv:1902.09907
  59. 59.
    Yang D L, Zhong X Y, Gu D B, Peng X F, Yang G L and Zou C S. 2020. Unsupervised learning of depth estimation, camera motion prediction and dynamic object localization from video. International Journal of Advanced Robotic Systems, 172): (172988142090965
  60. 60.
    Yin Z C and Shi J P. 2018. GeoNet: unsupervised learning of dense depth, optical flow and camera pose//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City: IEEE: 1983-1992
  61. 61.
    Zhang S, Zhang J and Tao D C. 2022. Towards scale-aware, robust, and generalizable unsupervised monocular depth estimation by integrating IMU motion dynamics//17th European Conference on Computer Vision. Tel Aviv: Springer: 143-160
  62. 62.
    Zhao C Q, Sun Q Y, Zhang C Z, Tang Y and Qian F. 2020. Monocular depth estimation based on deep learning: an overview. Science China Technological Sciences, 63(9): 1612-1627
  63. 63.
    Zhou T H, Brown M, Snavely N and Lowe D G. 2017. Unsupervised learning of depth and ego-motion from video//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu: IEEE: 6612-6619

Читать полностью

The above content is generated by Large Model Translation. The translated content is for reference only. We do not assume any commercial or legal responsibilty for any consequences arising from the use of our website