{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T09:15:52Z","timestamp":1785316552102,"version":"3.55.0"},"reference-count":81,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2020,1,27]],"date-time":"2020-01-27T00:00:00Z","timestamp":1580083200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Deanship of Scientific Research at King Saud University through the Local Research Group Program","award":["RG-1435-050"],"award-info":[{"award-number":["RG-1435-050"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Exploring the relevance between images and their respective natural language descriptions, due to its paramount importance, is regarded as the next frontier in the general computer vision literature. Thus, recently several works have attempted to map visual attributes onto their corresponding textual tenor with certain success. However, this line of research has not been widespread in the remote sensing community. On this point, our contribution is three-pronged. First, we construct a new dataset for text-image matching tasks, termed TextRS, by collecting images from four well-known different scene datasets, namely AID, Merced, PatternNet, and NWPU datasets. Each image is annotated by five different sentences. All the five sentences were allocated by five people to evidence the diversity. Second, we put forth a novel Deep Bidirectional Triplet Network (DBTN) for text to image matching. Unlike traditional remote sensing image-to-image retrieval, our paradigm seeks to carry out the retrieval by matching text to image representations. To achieve that, we propose to learn a bidirectional triplet network, which is composed of Long Short Term Memory network (LSTM) and pre-trained Convolutional Neural Networks (CNNs) based on (EfficientNet-B2, ResNet-50, Inception-v3, and VGG16). Third, we top the proposed architecture with an average fusion strategy to fuse the features pertaining to the five image sentences, which enables learning of more robust embedding. The performances of the method expressed in terms Recall@K representing the presence of the relevant image among the top K retrieved images to the query text shows promising results as it yields 17.20%, 51.39%, and 73.02% for K = 1, 5, and 10, respectively.<\/jats:p>","DOI":"10.3390\/rs12030405","type":"journal-article","created":{"date-parts":[[2020,1,27]],"date-time":"2020-01-27T11:41:57Z","timestamp":1580125317000},"page":"405","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":115,"title":["TextRS: Deep Bidirectional Triplet Network for Matching Text to Remote Sensing Images"],"prefix":"10.3390","volume":"12","author":[{"given":"Taghreed","family":"Abdullah","sequence":"first","affiliation":[{"name":"Department of Studies in Computer Science, University of Mysore, Manasagangothri, Mysore 570006, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9287-0596","authenticated-orcid":false,"given":"Yakoub","family":"Bazi","sequence":"additional","affiliation":[{"name":"Computer Engineering Department, College of Computer and Information Sciences, King Saud University, Riyadh 11543, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8105-9746","authenticated-orcid":false,"given":"Mohamad M.","family":"Al Rahhal","sequence":"additional","affiliation":[{"name":"Information System Department, College of Applied Computer Science, King Saud University, Riyadh 11543, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mohamed L.","family":"Mekhalfi","sequence":"additional","affiliation":[{"name":"Department of Information Engineering and Computer Science, University of Trento, Disi Via Sommarive 9, Povo, 28123 Trento, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lalitha","family":"Rangarajan","sequence":"additional","affiliation":[{"name":"Department of Studies in Computer Science, University of Mysore, Manasagangothri, Mysore 570006, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mansour","family":"Zuair","sequence":"additional","affiliation":[{"name":"Computer Engineering Department, College of Computer and Information Sciences, King Saud University, Riyadh 11543, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,1,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Al Rahhal, M.M., Bazi, Y., Abdullah, T., Mekhalfi, M.L., AlHichri, H., and Zuair, M. (2018). Learning a Multi-Branch Neural Network from Multiple Sources for Knowledge Adaptation in Remote Sensing Imagery. Remote Sens., 10.","DOI":"10.3390\/rs10121890"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"3023","DOI":"10.1109\/TGRS.2013.2268736","article-title":"Remote Sensing Image Retrieval With Global Morphological Texture Descriptors","volume":"52","author":"Aptoula","year":"2014","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"120","DOI":"10.1016\/j.isprsjprs.2017.11.021","article-title":"A new deep convolutional neural network for fast hyperspectral image classification","volume":"145","author":"Paoletti","year":"2018","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"2288","DOI":"10.1109\/36.868886","article-title":"Interactive learning and probabilistic retrieval in remote sensing image archives","volume":"38","author":"Schroder","year":"2000","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"606","DOI":"10.1109\/JSTSP.2011.2139193","article-title":"A Survey of Active Learning Algorithms for Supervised Remote Sensing Image Classification","volume":"5","author":"Tuia","year":"2011","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1865","DOI":"10.1109\/JPROC.2017.2675998","article-title":"Remote Sensing Image Scene Classification: Benchmark and State of the Art","volume":"105","author":"Cheng","year":"2017","journal-title":"Proc. IEEE"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2155","DOI":"10.1109\/LGRS.2015.2453130","article-title":"Land-Use Classification With Compressive Sensing Multifeature Fusion","volume":"12","author":"Mekhalfi","year":"2015","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Mekhalfi, M.L., and Melgani, F. (2015, January 26\u201331). Sparse modeling of the land use classification problem. Proceedings of the 2015 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Milan, Italy.","DOI":"10.1109\/IGARSS.2015.7326633"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"6281","DOI":"10.1080\/01431161.2018.1458346","article-title":"Land-use scene classification based on a CNN using a constrained extreme learning machine","volume":"39","author":"Weng","year":"2018","journal-title":"Int. J. Remote Sens."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1895","DOI":"10.1109\/LGRS.2016.2616440","article-title":"Deep Filter Banks for Land-Use Scene Classification","volume":"13","author":"Wu","year":"2016","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Shao, Z., Yang, K., and Zhou, W. (2018). Performance Evaluation of Single-Label and Multi-Label Remote Sensing Image Retrieval Using a Dense Labeling Dataset. Remote Sens., 10.","DOI":"10.3390\/rs10060964"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1144","DOI":"10.1109\/TGRS.2017.2760909","article-title":"Multi-label Remote Sensing Image Retrieval using a Semi-Supervised Graph-Theoretic Method","volume":"56","author":"Chaudhuri","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Shao, Z., Yang, K., and Zhou, W. (2018). Correction: Shao, Z.; et al. A Benchmark Dataset for Performance Evaluation of Multi-Label Remote Sensing Image Retrieval. Remote Sens., 10.","DOI":"10.3390\/rs10060964"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Bosilj, P., Aptoula, E., Lef\u00e8vre, S., and Kijak, E. (2016). Retrieval of Remote Sensing Images with Pattern Spectra Descriptors. ISPRS Int. J. Geo-Inf., 5.","DOI":"10.3390\/ijgi5120228"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"095994","DOI":"10.1117\/1.JRS.9.095994","article-title":"Dual-tree complex wavelet transform applied on color descriptors for remote-sensed images retrieval","volume":"9","author":"Sebai","year":"2015","journal-title":"J. Appl. Remote Sens."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Bouteldja, S., and Kourgli, A. (2015, January 10\u201312). Multiscale texture features for the retrieval of high resolution satellite images. Proceedings of the 2015 International Conference on Systems, Signals and Image Processing (IWSSIP), London, UK.","DOI":"10.1109\/IWSSIP.2015.7314204"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"083584","DOI":"10.1117\/1.JRS.8.083584","article-title":"Improved color texture descriptors for remote sensing image retrieval","volume":"8","author":"Shao","year":"2014","journal-title":"J. Appl. Remote Sens."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1603","DOI":"10.1109\/TGRS.2010.2088404","article-title":"Entropy-Balanced Bitmap Tree for Shape-Based Object Retrieval From Large-Scale Satellite Imagery Databases","volume":"49","author":"Scott","year":"2011","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive Image Features from Scale-Invariant Keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Bay, H., Tuytelaars, T., and Gool, L.V. (2006, January 7\u201313). SURF: Speeded Up Robust Features. Proceedings of the Computer Vision-ECCV 2006, Berlin, Heidelberg.","DOI":"10.1007\/11744023_32"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"273","DOI":"10.1080\/17538947.2014.882420","article-title":"An improved Bag-of-Words framework for remote sensing image retrieval in large-scale image databases","volume":"8","author":"Yang","year":"2015","journal-title":"Int. J. Digit. Earth"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"J\u00e9gou, H., Douze, M., Schmid, C., and P\u00e9rez, P. (2010, January 13\u201318). Aggregating local descriptors into a compact image representation. Proceedings of the 2010 IEEE computer society conference on computer vision and pattern recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5540039"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1793","DOI":"10.1109\/TGRS.2015.2488681","article-title":"Scene Classification via a Gradient Boosting Random Convolutional Network Framework","volume":"54","author":"Zhang","year":"2016","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1109\/MGRS.2016.2540798","article-title":"Deep Learning for Remote Sensing Data: A Technical Tutorial on the State of the Art","volume":"4","author":"Zhang","year":"2016","journal-title":"IEEE Geosci. Remote Sens. Mag."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1145\/3065386","article-title":"ImageNet Classification with Deep Convolutional Neural Networks","volume":"60","author":"Krizhevsky","year":"2017","journal-title":"Commun ACM"},{"key":"ref_27","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv, Available online: https:\/\/arxiv.org\/abs\/1409.1556."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"142","DOI":"10.1109\/TPAMI.2015.2437384","article-title":"Region-Based Convolutional Networks for Accurate Object Detection and Segmentation","volume":"38","author":"Girshick","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Noh, H., Hong, S., and Han, B. (2015, January 7\u201313). Learning Deconvolution Network for Semantic Segmentation. Proceedings of the IEEE international conference on computer vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.178"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1016\/j.isprsjprs.2017.11.004","article-title":"A semi-supervised generative framework with deep learning features for high-resolution remote sensing image scene classification","volume":"145","author":"Han","year":"2018","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"677","DOI":"10.1109\/TPAMI.2016.2599174","article-title":"Long-Term Recurrent Convolutional Networks for Visual Recognition and Description","volume":"39","author":"Donahue","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_34","unstructured":"Du, Y., Wang, W., and Wang, L. (2015, January 7\u201312). Hierarchical recurrent neural network for skeleton based action recognition. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA."},{"key":"ref_35","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R., and Bengio, Y. (2015, January 6\u201311). Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. Proceedings of the International Conference on Machine Learning (ICML), Lille, France."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"740","DOI":"10.1080\/2150704X.2019.1647368","article-title":"Enhancing remote sensing image retrieval using a triplet deep metric learning network","volume":"41","author":"Cao","year":"2020","journal-title":"Int. J. Remote Sens."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Tang, X., Zhang, X., Liu, F., and Jiao, L. (2018). Unsupervised Deep Feature Learning for Remote Sensing Image Retrieval. Remote Sens., 10.","DOI":"10.3390\/rs10081243"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"5288","DOI":"10.1109\/TIP.2018.2845136","article-title":"Dynamic Match Kernel With Deep Convolutional Features for Image Retrieval","volume":"27","author":"Yang","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"950","DOI":"10.1109\/TGRS.2017.2756911","article-title":"Large-Scale Remote Sensing Image Retrieval by Deep Hashing Neural Networks","volume":"56","author":"Li","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"464","DOI":"10.1109\/LGRS.2017.2651056","article-title":"Partial Randomness Hashing for Large-Scale Remote Sensing Image Retrieval","volume":"14","author":"Li","year":"2017","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"6521","DOI":"10.1109\/TGRS.2018.2839705","article-title":"Learning Source-Invariant Deep Hashing Convolutional Neural Networks for Cross-Source Remote Sensing Image Retrieval","volume":"56","author":"Li","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"987","DOI":"10.1109\/LGRS.2016.2558289","article-title":"Region-Based Retrieval of Remote Sensing Images Using an Unsupervised Graph-Theoretic Approach","volume":"13","author":"Chaudhuri","year":"2016","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_43","unstructured":"Zhou, W., Deng, X., and Shao, Z. (2018). Region Convolutional Features for Multi-Label Remote Sensing Image Retrieval. arXiv, Available online: https:\/\/arxiv.org\/abs\/1807.08634."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"2473","DOI":"10.1109\/JSTARS.2018.2832985","article-title":"A Novel System for Content-Based Retrieval of Single and Multi-Label High-Dimensional Remote Sensing Images","volume":"11","author":"Dai","year":"2018","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Wu, Q., Shen, C., Liu, L., Dick, A., and Hengel, A.v.d. (2016, January 27\u201330). What Value Do Explicit High Level Concepts Have in Vision to Language Problems?. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.29"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014). Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. arXiv, Available online: https:\/\/arxiv.org\/abs\/1406.1078.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"652","DOI":"10.1109\/TPAMI.2016.2587640","article-title":"Show and Tell: Lessons Learned from the 2015 MSCOCO Image Captioning Challenge","volume":"39","author":"Vinyals","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_48","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012). Imagenet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, MIT Press."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Fang, H., Gupta, S., Iandola, F.N., Srivastava, R.K., Deng, L., Dollar, P., Gao, J., He, X., Mitchell, M., and Platt, J. (2015, January 7\u201312). From captions to visual concepts and back. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298754"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"You, Q., Jin, H., Wang, Z., Fang, C., and Luo, J. (2016, January 27\u201330). Image Captioning with Semantic Attention. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vega, NV, USA.","DOI":"10.1109\/CVPR.2016.503"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Chen, L., Zhang, H., Xiao, J., Nie, L., Shao, J., Liu, W., and Chua, T.S. (2017, January 21\u201326). Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.667"},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"2008","DOI":"10.1109\/TIP.2018.2882225","article-title":"Bi-directional Spatial-Semantic Attention Networks for Image-Text Matching","volume":"28","author":"Huang","year":"2018","journal-title":"IEEE Trans. Image Process. Publ. IEEE Signal Process. Soc."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"394","DOI":"10.1109\/TPAMI.2018.2797921","article-title":"Learning Two-Branch Neural Networks for Image-Text Matching Tasks","volume":"41","author":"Wang","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Zhang, Y., and Lu, H. (2018, January 8\u201314). Deep Cross-Modal Projection Learning for Image-Text Matching. Proceedings of the European Conference on Computer Vision\u2014ECCV 2018, Munich, Germany.","DOI":"10.1007\/978-3-030-01246-5_42"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Yao, T., Pan, Y., Li, Y., Qiu, Z., and Mei, T. (2017, January 22\u201329). Boosting image captioning with attributes. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.524"},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"2183","DOI":"10.1109\/TGRS.2017.2776321","article-title":"Exploring Models and Data for Remote Sensing Image Caption Generation","volume":"56","author":"Lu","year":"2018","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"3623","DOI":"10.1109\/TGRS.2017.2677464","article-title":"Can a Machine Generate Humanlike Language Descriptions for a Remote Sensing Image?","volume":"55","author":"Shi","year":"2017","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Zhang, X., Wang, X., Tang, X., Zhou, H., and Li, C. (2019). Description Generation for Remote Sensing Images Using Attribute Attention Mechanism. Remote Sens., 11.","DOI":"10.3390\/rs11060612"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201322). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 27\u201330). Rethinking the Inception Architecture for Computer Vision. Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_61","unstructured":"Tan, M., and Le, Q.V. (2019). EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. arXiv, Available online: https:\/\/arxiv.org\/abs\/1905.11946."},{"key":"ref_62","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long Short-Term Memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_63","first-page":"207","article-title":"Distance Metric Learning for Large Margin Nearest Neighbor Classification","volume":"10","author":"Weinberger","year":"2009","journal-title":"J. Mach. Learn. Res."},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"Wang, J., Song, Y., Leung, T., and Rosenberg, C. (2014, January 23\u201328). Learning Fine-Grained Image Similarity with Deep Ranking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.180"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Feragen, A., Pelillo, M., and Loog, M. (2015, January 12\u201314). Deep Metric Learning Using Triplet Network. Proceedings of the Similarity-Based Pattern Recognition, Copenhagen, Denmark.","DOI":"10.1007\/978-3-319-24261-3"},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Law, M.T., Thome, N., and Cord, M. (2013, January 1\u20138). Quadruplet-Wise Image Similarity Learning. Proceedings of the 2013 IEEE International Conference on Computer Vision, Sydney, NSW, Australia.","DOI":"10.1109\/ICCV.2013.38"},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Oh Song, H., Xiang, Y., Jegelka, S., and Savarese, S. (2016, January 27\u201330). Deep Metric Learning via Lifted Structured Feature Embedding. Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.434"},{"key":"ref_68","unstructured":"Sohn, K. (2016, January 5\u201310). Improved deep metric learning with multi-class n-pair loss objective. Proceedings of the Advances in Neural Information Processing Systems (NIPS), Barcelona, Spain."},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Wang, J., Zhou, F., Wen, S., Liu, X., and Lin, Y. (2017, January 22\u201329). Deep Metric Learning with Angular Loss. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.283"},{"key":"ref_70","doi-asserted-by":"crossref","unstructured":"Huang, J., Feris, R., Chen, Q., and Yan, S. (2015, January 7\u201313). Cross-Domain Image Retrieval with a Dual Attribute-Aware Ranking Network. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.127"},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Lai, H., Pan, Y., Liu, Y., and Yan, S. (2015, January 7\u201312). Simultaneous Feature Learning and Hash Coding With Deep Neural Networks. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298947"},{"key":"ref_72","doi-asserted-by":"crossref","unstructured":"Zhuang, B., Lin, G., Shen, C., and Reid, I. (2016, January 27\u201330). Fast Training of Triplet-Based Deep Binary Embedding Networks. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.641"},{"key":"ref_73","doi-asserted-by":"crossref","unstructured":"Gordo, A., Almazan, J., Revaud, J., and Larlus, D. (2016, January 11\u201314). Deep image retrieval: Learning global representations for image search. VI. Proceedings of the Computer Vision\u2014ECCV 2016, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46466-4_15"},{"key":"ref_74","doi-asserted-by":"crossref","unstructured":"Yuan, Y., Yang, K., and Zhang, C. (2017, January 22\u201329). Hard-Aware Deeply Cascaded Embedding. Proceedings of the IEEE international conference on computer vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.94"},{"key":"ref_75","unstructured":"Parkhi, O.M., Vedaldi, A., and Zisserman, A. Deep Face Recognition. Proceedings of the British Machine Vision Conference (BMVC)."},{"key":"ref_76","doi-asserted-by":"crossref","unstructured":"Wang, L., Li, Y., and Lazebnik, S. (2016, January 27\u201330). Learning Deep Structure-Preserving Image-Text Embeddings. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.541"},{"key":"ref_77","doi-asserted-by":"crossref","unstructured":"Schroff, F., Kalenichenko, D., and Philbin, J. (2015, January 7\u201312). FaceNet: A Unified Embedding for Face Recognition and Clustering. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"ref_78","doi-asserted-by":"crossref","unstructured":"Harwood, B., VijayKumar, B.G., Carneiro, G., Reid, I., and Drummond, T. (2017, January 22\u201329). Smart Mining for Deep Metric Learning. Proceedings of the IEEE international conference on computer vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.307"},{"key":"ref_79","doi-asserted-by":"crossref","unstructured":"Wu, C.-Y., Manmatha, R., Smola, A.J., and Kr\u00e4henb\u00fchl, P. (2017, January 22\u201329). Sampling Matters in Deep Embedding Learning. Proceedings of the IEEE international conference on computer vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.309"},{"key":"ref_80","doi-asserted-by":"crossref","unstructured":"Ge, W., Huang, W., Dong, D., and Scott, M.R. (2018, January 8\u201314). Deep Metric Learning with Hierarchical Triplet Loss. Proceedings of the Computer Vision\u2014ECCV 2018, Munich, Germany.","DOI":"10.1007\/978-3-030-01231-1_17"},{"key":"ref_81","doi-asserted-by":"crossref","unstructured":"Plummer, B.A., Wang, L., Cervantes, C.M., Caicedo, J.C., Hockenmaier, J., and Lazebnik, S. (2015, January 13\u201316). Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.303"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/3\/405\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T14:08:41Z","timestamp":1760364521000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/12\/3\/405"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,1,27]]},"references-count":81,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2020,2]]}},"alternative-id":["rs12030405"],"URL":"https:\/\/doi.org\/10.3390\/rs12030405","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,1,27]]}}}