{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T16:53:19Z","timestamp":1777654399419,"version":"3.51.4"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,10,23]],"date-time":"2023-10-23T00:00:00Z","timestamp":1698019200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272018"],"award-info":[{"award-number":["62272018"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,3,31]]},"abstract":"<jats:p>\n            Robust and accurate camera pose estimation is fundamental in computer vision. Learning-based regression approaches acquire six-degree-of-freedom camera parameters accurately from visual cues of an input image. However, most are trained on street-view and landmark datasets. These approaches can hardly be generalized to overlooking use cases, such as the calibration of the surveillance camera and unmanned aerial vehicle. Besides, reference images captured from the real world are rare and expensive, and their diversity is not guaranteed. In this article, we address the problem of using alternative virtual images for visual localization training. This work has the following principle contributions: First, we present a new challenging localization dataset containing six reconstructed large-scale three-dimensional scenes, 10,594 calibrated photographs with condition changes, and 300k virtual images with pixelwise labeled depth, relative surface normal, and semantic segmentation. Second, we present a flexible multi-feature fusion network trained on virtual image datasets for robust image retrieval. Third, we propose an end-to-end confidence map prediction network for feature filtering and pose estimation. We demonstrate that large-scale rendered virtual images are beneficial to visual localization. Using virtual images can solve the diversity problem of real images and leverage labeled multi-feature data for deep learning. Experimental results show that our method achieves remarkable performance surpassing state-of-the-art approaches. To foster research on improvement for visual localization using synthetic images, we release our benchmark at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/YuanXiong\/contributions\">https:\/\/github.com\/YuanXiong\/contributions<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3622788","type":"journal-article","created":{"date-parts":[[2023,9,4]],"date-time":"2023-09-04T13:17:40Z","timestamp":1693833460000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["VirtualLoc: Large-scale Visual Localization Using Virtual Images"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7253-4998","authenticated-orcid":false,"given":"Yuan","family":"Xiong","sequence":"first","affiliation":[{"name":"State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7896-8765","authenticated-orcid":false,"given":"Jingru","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5825-7517","authenticated-orcid":false,"given":"Zhong","family":"Zhou","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, P.R. China and Zhongguancun Laboratory, P.R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,10,23]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/2422956.2422964"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.572"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/IVS.2014.6856605"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01529"},{"issue":"9","key":"e_1_3_1_6_2","first-page":"5847","article-title":"Visual camera re-localization from RGB and RGB-D images using DSAC","volume":"44","author":"Brachmann Eric","year":"2021","unstructured":"Eric Brachmann and Carsten Rother. 2021. Visual camera re-localization from RGB and RGB-D images using DSAC. IEEE Trans. Pattern Anal. Mach. Intell. 44, 9 (2021), 5847\u20135865.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00277"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"e_1_3_1_9_2","unstructured":"Weifeng Chen Zhao Fu Dawei Yang and Jia Deng. 2016. Single-image depth perception in the wild. Adv. Neural Inf. Process. Syst. 29 (2016) 730\u2013738."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.350"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS51168.2021.9636156"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3167919"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364919863090"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00600"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00828"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548100"},{"key":"e_1_3_1_17_2","unstructured":"Zan Gao Chao Sun Zhiyong Cheng Weili Guan Anan Liu and Meng Wang. 2023. TBNet: A two-stream boundary aware network for generic image manipulation localization. IEEE Trans. Knowl. Data Eng. 35 7 (2023) 7541\u20137556."},{"key":"e_1_3_1_18_2","article-title":"S2dnet: Learning accurate correspondences for sparse-to-dense feature matching","author":"Germain Hugo","year":"2020","unstructured":"Hugo Germain, Guillaume Bourmaud, and Vincent Lepetit. 2020. S2dnet: Learning accurate correspondences for sparse-to-dense feature matching. arXiv:2004.01673. Retrieved from https:\/\/arxiv.org\/abs\/2004.01673","journal-title":"arXiv:2004.01673"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-017-1016-8"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3424115"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01392"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2926463"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3136714"},{"key":"e_1_3_1_24_2","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV\u201912)","author":"J\u00e9gou Herv\u00e9","year":"2012","unstructured":"Herv\u00e9 J\u00e9gou and Ondrej Chum. 2012. Negative evidences and co-occurrences in image retrieval: The benefit of PCA and whitening. In Proceedings of the European Conference on Computer Vision (ECCV\u201912)."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.336"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00012"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_15"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000029664.99615.94"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298925"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/2406367.2406372"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20047-2_34"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01266"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01476"},{"key":"e_1_3_1_34_2","first-page":"652","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Qi Charles R.","year":"2017","unstructured":"Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. 2017. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 652\u2013660."},{"key":"e_1_3_1_35_2","unstructured":"Charles Ruizhongtai Qi Li Yi Hao Su and Leonidas J Guibas. 2017. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Adv. Neural Inf. Process. Syst. 30 (2017) 5105\u20135114."},{"key":"e_1_3_1_36_2","first-page":"12414","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"Revaud Jerome","year":"2019","unstructured":"Jerome Revaud, Philippe Weinzaepfel, C\u00e9sar De Souza, and Martin Humenberger. 2019. R2D2: Repeatable and reliable detector and descriptor. In Proceedings of the 33rd International Conference on Neural Information Processing Systems. 12414\u201312424."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01300"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00499"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00326"},{"key":"e_1_3_1_40_2","doi-asserted-by":"crossref","unstructured":"Linus Sv\u00e4rm Olof Enqvist Fredrik Kahl and Magnus Oskarsson. 2017. City-scale localization for cameras with known vertical direction. IEEE Trans. Pattern Anal. Mach. Intell. 39 7 (2017) 1455\u20131461.","DOI":"10.1109\/TPAMI.2016.2598331"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2611662"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00897"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00342"},{"key":"e_1_3_1_44_2","doi-asserted-by":"crossref","first-page":"835","DOI":"10.1145\/1179352.1141964","volume-title":"ACM Siggraph 2006 Papers","author":"Snavely Noah","year":"2006","unstructured":"Noah Snavely, Steven M. Seitz, and Richard Szeliski. 2006. Photo tourism: Exploring photo collections in 3D. In ACM Siggraph 2006 Papers. 835\u2013846."},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","unstructured":"Torsten Sattler Akihiko Torii Josef Sivic Marc Pollefeys Hajime Taira Masatoshi Okutomi and Tomas Pajdla. 2017. Are large-scale 3D models really necessary for accurate visual localization? In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1637\u20131646.","DOI":"10.1109\/CVPR.2017.654"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8463150"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3032010"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2017.83"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01216-8_24"},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","unstructured":"Will Maddern Geoffrey Pascoe Chris Linegar and Paul Newman. 2017. 1 year 1000 km: The Oxford RobotCar dataset. The International Journal of Robotics Research . 36 1 (2017) 3\u201315.","DOI":"10.1177\/0278364916679498"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58517-4_28"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298790"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00470"},{"key":"e_1_3_1_54_2","first-page":"1","article-title":"An efficient 3-D point cloud place recognition approach based on feature point extraction and transformer","volume":"71","author":"Ye Tao","year":"2022","unstructured":"Tao Ye, Xiangming Yan, Shouan Wang, Yunwang Li, and Fuqiang Zhou. 2022. An efficient 3-D point cloud place recognition approach based on feature point extraction and transformer. IEEE Trans. Instrum. Meas. 71 (2022), 1\u20139.","journal-title":"IEEE Trans. Instrum. Meas."},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2019.2898737"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2871832"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3622788","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3622788","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:04Z","timestamp":1750178224000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3622788"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,23]]},"references-count":55,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,3,31]]}},"alternative-id":["10.1145\/3622788"],"URL":"https:\/\/doi.org\/10.1145\/3622788","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,23]]},"assertion":[{"value":"2023-05-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-08-27","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}