Cross-view scene image localization with Triplet Network integrating NetVLAD and Fully Connected Layers

  • role: First author第一作者
  • Affiliation:

    School of Earth Sciences and Engineering, Hohai University, Nanjing 211100, China

  • Email:zhaohui.xue@hhu.edu.cn
  • Introduction:1984,,,E-mail zhaohui.xue@hhu.edu.cn
XUE Zhaohui1,  
  • Affiliation:

    School of Earth Sciences and Engineering, Hohai University, Nanjing 211100, China

ZHOU Yiyang1,  
  • Affiliation:

    School of Computer Science and Technology, University of Science and Technology of China, Hefei 230026, China

QIANG Yonggang2,  
  • Affiliation:

    National Engineering Laboratory for Social Security Risk Perception and Prevention and Control of Big Data Application, China Academy of Electronics, Beijing 100041, China

LIU Yifeng3,  
  • Affiliation:

    National Engineering Laboratory for Social Security Risk Perception and Prevention and Control of Big Data Application, China Academy of Electronics, Beijing 100041, China

LIN Hui3

resumen

Cross-view scene image matching and positioning have a wide range of applications in target search, combating crime, and positioning. With the development of deep learning, neural networks have played an important role in this issue. Given the problem of cross-view scene image matching and positioning between street view and bird’s eye images, the neural network model’s convergence is slow, and the feature correlation is weak. This paper proposes a triplet network model (Tri-NetVLAD) that combines NetVLAD and a fully connected layer and improves DBL Loss (ADBL loss). The proposed method can not only improve the convergence speed and stability of the network but also the overall positioning accuracy of the model.The proposed Tri-NetVLAD model extracts the local features of the three input images through a triplet network and inputs the local features to the fully connected and NetVLAD layers to obtain the feature vector and the global feature descriptor. The global feature descriptor can obtain the relative distribution between features, and on this basis, incorporate feature vectors, which can preserve the differences between features to improve the positioning accuracy of the model. ADBL loss improves the model’s ability to discriminate difficult samples by introducing parameters and the positioning accuracy of the model.The proposed Tri-NetVLAD is compared with several existing methods, namely, MCVPlaces, Triplet eDBL-Net, and CVM-Net, and loss functions, namely, contrastive loss, triplet loss, and DBL loss. In the US vo and hays dataset, the highest positioning accuracy of 63.5% is achieved, proving that the triplet network that combines the NetVLAD and fully connected layers can effectively improve the positioning accuracy with the ADBL Loss.Compared with existing methods, the proposed Tri-NetVLAD has the following advantages. (1) The Triplet network can increase the Euclidean distance between unmatched images while reducing the Euclidean distance between matched images. (2) The introduction of NetVLAD can aggregate the local features extracted by CNN to obtain global feature descriptors and the distribution relationship between features. (3) The fusing of the Fully Connected Layer adds the feature vector obtained through the fully connected layer to the global feature descriptor, so that the final feature vector not only represents the distribution relationship between features, but also retains the differences between features. (4) The improved loss function ADBL Loss can accelerate the gradient convergence speed and improve the overall positioning accuracy.

palabra clave

cross-view;scene image matching and geolocation;Triplet Network;NetVLAD;CNN

References

  1. 1.
    Altwaijry H, Trulls E, Hays J, Fua P and Belongie S. 2016. Learning to Match Aerial Images with Deep Attentive Architectures//Proceedings of 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas, NV, USA: IEEE: 3539-3547
  2. 2.
    Arandjelović R, Gronat P, Torii A, Pajdla T and Sivic J. 2018. NetVLAD: CNN architecture for weakly supervised place recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(6): 1437-1451
  3. 3.
    Chen W H, Chen X T, Zhang J G and Huang K Q. 2017. Beyond Triplet Loss: a deep quadruplet network for person Re-identification//Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu, HI, USA: IEEE: 403-412
  4. 4.
    Hadsell R, Chopra S and LeCun Y. 2006. Dimensionality reduction by learning an invariant mapping//Proceedings of 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06). New York, NY, USA: IEEE: 1735-1742
  5. 5.
    Hammoud R I, Kuzdeba S A, Berard B, Tom V, Ivey R, Bostwick R, HandUber J, Vinciguerra L, Shnidman N and Smiley B. 2013. Overhead-based image and video geo-localization framework//Proceedings of 2013 IEEE Conference on Computer Vision and Pattern Recognition Workshops. Portland, OR, USA: IEEE: 320-327
  6. 6.
    Hu S X, Feng M D, Nguyen R M H and Lee G H. 2018. CVM-Net: cross-view matching network for image-based ground-to-aerial geo-localization//Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City, UT, USA: IEEE: 7258-7267
  7. 7.
    Jégou H, Perronnin F, Douze M, Sánchez J, Pérez P and Schmid C. 2012. Aggregating local image descriptors into compact codes. IEEE Transactions on Pattern Analysis and Machine Intelligence, 34(9): 1704-1716
  8. 8.
    Kingma D P and Ba J. 2014. Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980
  9. 9.
    Krizhevsky A, Sutskever I and Hinton G E. 2012. Imagenet classification with deep convolutional neural networks//Proceedings of the 25th International Conference on Neural Information Processing Systems. Lake Tahoe, Nevada, USA: ACM: 1097-1105
  10. 10.
    Lin T Y, Belongie S and Hays J. 2013. Cross-view image geolocalization//Proceedings of 2013 IEEE Conference on Computer Vision and Pattern Recognition. Portland, OR, USA: IEEE: 891-898
  11. 11.
    Lin T Y, Cui Y, Belongie S and Hays J. 2015. Learning deep representations for ground-to-aerial geolocalization//Proceedings of 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Boston, MA, USA: IEEE: 5007-5015
  12. 12.
    Schindler G, Brown M and Szeliski R. 2007. City-scale location recognition//Proceedings of 2007 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Minneapolis, MN, USA: IEEE: 1-7
  13. 13.
    Tian Y C, Chen C and Shah M. 2017. Cross-view image matching for geo-localization in urban environments//Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu, HI, USA: IEEE: 1998-2006
  14. 14.
    Vo N N and Hays J. 2016. Localizing and orienting street views using overhead imagery//Proceedings of the 14th European Conference on Computer Vision (ECCV). Amsterdam, The Netherlands: Springer: 494-509
  15. 15.
    Workman S, Souvenir R and Jacobs N. 2015. Wide-area image geolocalization with aerial reference imagery//Proceedings of 2015 IEEE International Conference on Computer Vision (ICCV). Santiago, Chile: IEEE: 3961-3969
  16. 16.
    Zhang H Q, Liu X Y, Yang S and Li Y. 2017. Retrieval of remote sensing images based on semisupervised deep learning. Journal of Remote Sensing, 21(3): 406-414
  17. 17.
    Zhang L and Liao M S. 2006. A context aware fuzzy clustering method for remote sensing images. Journal of Remote Sensing, 2006 (01): 58-65
  18. 18.
    Zhao L J and Tang P. 2016. Scalability analysis of typical remote sensing data classification methods: a case of remote sensing image scene. Journal of Remote Sensing, 20(2): 157-171

Leer el texto completo

The above content is generated by Large Model Translation. The translated content is for reference only. We do not assume any commercial or legal responsibilty for any consequences arising from the use of our website