{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T05:05:27Z","timestamp":1783400727029,"version":"3.54.6"},"reference-count":191,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2022,11,21]],"date-time":"2022-11-21T00:00:00Z","timestamp":1668988800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004263","name":"Funda\u00e7\u00e3o de Amparo \u00e0 Pesquisa do Estado do Rio Grande do Sul","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004263","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100003593","name":"Conselho Nacional de Desenvolvimento Cient\u00edfico and Tecnol\u00f3gico","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100003593","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Coordena\u00e7\u00e3o de Aperfei\u00e7oamento de Pessoal de N\u00edvel Superior (CAPES), Brazil"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2023,4,30]]},"abstract":"<jats:p>This article provides a comprehensive survey on pioneer and state-of-the-art 3D scene geometry estimation methodologies based on single, two, or multiple images captured under omnidirectional optics. We first revisit the basic concepts of the spherical camera model and review the most common acquisition technologies and representation formats suitable for omnidirectional (also called 360\u00b0, spherical or panoramic) images and videos. We then survey monocular layout and depth inference approaches, highlighting the recent advances in learning-based solutions suited for spherical data. The classical stereo matching is then revised on the spherical domain, where methodologies for detecting and describing sparse and dense features become crucial. The stereo matching concepts are then extrapolated for multiple view camera setups, categorizing them among light fields, multi-view stereo, and structure from motion (or visual simultaneous localization and mapping). We also compile and discuss commonly adopted datasets and figures of merit indicated for each purpose and list recent results for completeness. We conclude this article by pointing out current and future trends.<\/jats:p>","DOI":"10.1145\/3519021","type":"journal-article","created":{"date-parts":[[2022,3,4]],"date-time":"2022-03-04T12:29:59Z","timestamp":1646396999000},"page":"1-39","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":34,"title":["3D Scene Geometry Estimation from 360\u00b0\u00a0Imagery: A Survey"],"prefix":"10.1145","volume":"55","author":[{"given":"Thiago L. T.","family":"da Silveira","sequence":"first","affiliation":[{"name":"Institute of Informatics, Federal University of Rio Grande do Sul (UFRGS), Porto Alegre, Brazil"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Paulo G. L.","family":"Pinto","sequence":"additional","affiliation":[{"name":"Institute of Informatics, Federal University of Rio Grande do Sul (UFRGS), Porto Alegre, Brazil"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jeffri","family":"Murrugarra-Llerena","sequence":"additional","affiliation":[{"name":"Institute of Informatics, Federal University of Rio Grande do Sul (UFRGS), Porto Alegre, Brazil"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cl\u00e1udio R.","family":"Jung","sequence":"additional","affiliation":[{"name":"Institute of Informatics, Federal University of Rio Grande do Sul (UFRGS), Porto Alegre, Brazil"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,11,21]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.120"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2016.2645119"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/2001269.2001293"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.408"},{"key":"e_1_3_2_6_2","first-page":"29","article-title":"Two-and three-view geometry for spherical cameras","volume":"105","author":"Akihiko T.","year":"2005","unstructured":"T. Akihiko, I. Atsushi, and N. Ohnishi. 2005. Two-and three-view geometry for spherical cameras. Workshop on Omnidirectional Vision, Camera Networks and Non-classical Cameras 105 (2005), 29\u201334.","journal-title":"Workshop on Omnidirectional Vision, Camera Networks and Non-classical Cameras"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33783-3_16"},{"key":"e_1_3_2_8_2","first-page":"13.1\u201313.11","article-title":"Fast explicit diffusion for accelerated features in nonlinear scale spaces","author":"Alcantarilla P.","year":"2013","unstructured":"P. Alcantarilla, J. Nuevo, and A. Bartoli. 2013. Fast explicit diffusion for accelerated features in nonlinear scale spaces. BMVC (2013), 13.1\u201313.11.","journal-title":"BMVC"},{"key":"e_1_3_2_9_2","first-page":"1","article-title":"A phase-based framework for optical flow estimation on omnidirectional images","volume":"10","author":"Alibouch B.","year":"2014","unstructured":"B. Alibouch, A. Radgui, C. Demonceaux, M. Rziza, and D. Aboutajdine. 2014. A phase-based framework for optical flow estimation on omnidirectional images. Signal Image Video Process 10 (2014), 1\u20138.","journal-title":"Signal Image Video Process"},{"issue":"1312","key":"e_1_3_2_10_2","first-page":"978","article-title":"Jump: Virtual reality video","volume":"3516","author":"Anderson R.","year":"2016","unstructured":"R. Anderson, D. Gallup, J. T. Barron, Janne K., N. Snavely, C. Hern\u00e1ndez, S. Agarwal, and S. M. Seitz. 2016. Jump: Virtual reality video. ACM Transactions on Graphics Article 3516, 1312 (2016), 978\u20131.","journal-title":"ACM Transactions on Graphics Article"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/AVSS.2007.4425344"},{"key":"e_1_3_2_12_2","unstructured":"I. Armeni S. Sax A. R. Zamir and S. Savarese. 2017. Joint 2D-3D-semantic data for indoor scene understanding. arxiv:1702.01105. Retrieved from https:\/\/arxiv.org\/abs\/1702.01105."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR48806.2021.9412745"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2009.5414552"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10851-011-0267-1"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.5194\/isprs-archives-XLII-2-W3-85-2017"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299076"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/11744023_32"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1175\/1520-0493(1999)127<2733:CTOASU>2.0.CO;2"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/SIBGRAPI54419.2021.00015"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2019.2898799"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417770"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/34.121791"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR48806.2021.9412035"},{"key":"e_1_3_2_25_2","article-title":"Monocular depth estimation: A survey","author":"Bhoi A.","year":"2019","unstructured":"A. Bhoi. 2019. Monocular depth estimation: A survey. CoRR (2019).","journal-title":"CoRR"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.12289"},{"key":"e_1_3_2_27_2","volume-title":"Blender - A 3D Modelling and Rendering Package","author":"Community Blender Online","year":"2020","unstructured":"Blender Online Community. 2020. Blender - A 3D Modelling and Rendering Package. Blender Foundation."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF00128525"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007928406666"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.546"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2011.222"},{"key":"e_1_3_2_32_2","first-page":"141","volume-title":"Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Caruso D.","year":"2015","unstructured":"D. Caruso, J. Engel, and D. Cremers. 2015. Large-scale direct SLAM for omnidirectional cameras. In Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems. 141\u2013148."},{"key":"e_1_3_2_33_2","article-title":"Matterport3D: Learning from RGB-D data in indoor environments","author":"Chang A.","year":"2017","unstructured":"A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niessner, M. Savva, S. Song, A. Zeng, and Y. Zhang. 2017. Matterport3D: Learning from RGB-D data in indoor environments. InProceedings of the International Conference on 3D Vision (2017).","journal-title":"Proceedings of the International Conference on 3D Vision"},{"key":"e_1_3_2_34_2","article-title":"Spherical CNNs","author":"Cohen T. S.","year":"2018","unstructured":"T. S. Cohen, M. Geiger, J. K\u00f6hler, and M. Welling. 2018. Spherical CNNs. InProceedings of the ICLR.","journal-title":"Proceedings of the ICLR"},{"key":"e_1_3_2_35_2","first-page":"525","article-title":"SphereNet: Learning spherical representations for detection and classification in omnidirectional images","author":"Coors B.","year":"2018","unstructured":"B. Coors, A. P. Condurache, and A. Geiger. 2018. SphereNet: Learning spherical representations for detection and classification in omnidirectional images. In Proceedings of the European Conference on Computer Vision.525\u2013541.","journal-title":"Proceedings of the European Conference on Computer Vision."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-011-0505-4"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.sigpro.2021.108277"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2018.8451769"},{"key":"e_1_3_2_39_2","first-page":"374","article-title":"Evaluation of keypoint extraction and matching for pose estimation using pairs of spherical images","author":"Silveira T. L. T. da","year":"2017","unstructured":"T. L. T. da Silveira and C. R. Jung. 2017. Evaluation of keypoint extraction and matching for pose estimation using pairs of spherical images. In Proceedings of the 30th SIBGRAPI Conference on Graphics, Patterns and Images. 374\u2013381.","journal-title":"Proceedings of the 30th SIBGRAPI Conference on Graphics, Patterns and Images"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/VR.2019.8798281"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01203"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.89"},{"key":"e_1_3_2_43_2","volume-title":"Proceedings of the ICLR","author":"Defferrard M.","year":"2020","unstructured":"M. Defferrard, M. Milani, F. Gusset, and N. Perraudin. 2020. DeepSphere: A graph-based spherical CNN. In Proceedings of the ICLR."},{"key":"e_1_3_2_44_2","volume-title":"Proceedings of the 7th World Congress on Intelligent Control and Automation.","author":"Deng X.","year":"2008","unstructured":"X. Deng, F. Wu, Y. Wu, and C. Wan. 2008. Automatic spherical panorama generation with two fisheye images. In Proceedings of the 7th World Congress on Intelligent Control and Automation."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2018.00060"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.316"},{"key":"e_1_3_2_47_2","unstructured":"M. Eder and J.-M. Frahm. 2019. Convolutions on spherical images. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops ."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2019.00018"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01244"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.304"},{"key":"e_1_3_2_51_2","first-page":"2366","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Eigen D.","year":"2014","unstructured":"D. Eigen, C. Puhrsch, and R. Fergus. 2014. Depth map prediction from a single image using a multi-scale deep network. In Proceedings of the Advances in Neural Information Processing Systems. 2366\u20132374."},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01261-8_4"},{"issue":"2","key":"e_1_3_2_53_2","first-page":"331","article-title":"Improving spherical photogrammetry using 360 \\( ^\\circ \\)  OMNI-Cameras: Use cases and new applications","volume":"42","author":"Fangi G.","year":"2018","unstructured":"G. Fangi, R. Pierdicca, M. Sturari, and E. S. Malinverni. 2018. Improving spherical photogrammetry using 360 \\( ^\\circ \\) OMNI-Cameras: Use cases and new applications. IISPRS Archives 42, 2 (2018), 331\u2013337.","journal-title":"IISPRS Archives"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000022288.19776.77"},{"key":"e_1_3_2_55_2","first-page":"1","article-title":"Corners for layout: End-to-end layout recovery from 360 images","author":"Fernandez-Labrador C.","year":"2020","unstructured":"C. Fernandez-Labrador, J. M. Facil, A. Perez-Yus, C. Demonceaux, J. Civera, and J. Guerrero. 2020. Corners for layout: End-to-end layout recovery from 360 images. IEEE Robotics and Automation Letters (2020), 1\u20131.","journal-title":"IEEE Robotics and Automation Letters"},{"key":"e_1_3_2_56_2","unstructured":"C. Fernandez-Labrador J. M. Facil A. Perez-Yus C. Demonceaux and J. J. Guerrero. 2018. PanoRoom: From the sphere to the 3D layout. (2018) 1\u20136. arxiv:1808.09879. Retrieved from https:\/\/arxiv.org\/abs\/1808.09879."},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2018.2850532"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2017.106"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2016.7899892"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1561\/0600000052"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2007.383246"},{"key":"e_1_3_2_62_2","first-page":"812","article-title":"Eliminating the blind spot: Adapting 3D object detection and monocular depth estimation to 360 \\( ^\\circ \\)  panoramic imagery","author":"Garanderie G. P. de La","year":"2018","unstructured":"G. P. de La Garanderie, A. A. Abarghouei, and T. P. Breckon. 2018. Eliminating the blind spot: Adapting 3D object detection and monocular depth estimation to 360 \\( ^\\circ \\) panoramic imagery. In Proceedings of the European Conference on Computer Vision (2018), 812\u2013830.","journal-title":"Proceedings of the European Conference on Computer Vision"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964964"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2018.8453486"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0097-8493(03)00038-4"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.519"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2621662"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-41914-0_40"},{"key":"e_1_3_2_69_2","first-page":"427","volume-title":"Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (Cat. No. 98CB36231).","author":"Shum H.-Y.","year":"1998","unstructured":"H.-Y. Shum, M. Han, and R. Szeliski. 1998. Interactive construction of 3D models from panoramic mosaics. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (Cat. No. 98CB36231).427\u2013433."},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1155\/2016\/8742920"},{"key":"e_1_3_2_71_2","first-page":"23.1\u201323.6","article-title":"A combined corner and edge detector","author":"Harris C.","year":"1988","unstructured":"C. Harris and M. Stephens. 1988. A combined corner and edge detector. Alvey Vision Conference (1988), 23.1\u201323.6.","journal-title":"Alvey Vision Conference"},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/34.601246"},{"key":"e_1_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.5555\/861369"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(81)90024-2"},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00936"},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46487-9_10"},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.1109\/VR.2017.7892229"},{"key":"e_1_3_2_78_2","first-page":"2695","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Xiao J.","year":"2012","unstructured":"J. Xiao, K. A. Ehinger, A. Oliva, and A. Torralba. 2012. Recognizing scene viewpoint using panoramic place representation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.2695\u20132702."},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2015.7139873"},{"key":"e_1_3_2_80_2","volume-title":"Proceedings of the International Conference on Robotics and Automation.","author":"Jiang C. \u201cM.\u201d","year":"2019","unstructured":"C. \u201cM.\u201d Jiang, J. Huang, K. Kashinath, Prabhat, P. Marcus, and M. Niessner. 2019. Spherical CNNs on unstructured grids. In Proceedings of the International Conference on Robotics and Automation."},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2021.3058957"},{"key":"e_1_3_2_82_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00097"},{"key":"e_1_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2017.226"},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.3390\/ijgi9050330"},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2007.4409198"},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.3390\/s20082272"},{"key":"e_1_3_2_87_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8682203"},{"key":"e_1_3_2_88_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2016.83"},{"key":"e_1_3_2_89_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2009.5457429"},{"key":"e_1_3_2_90_2","first-page":"169","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Kim H.","year":"2010","unstructured":"H. Kim and A. Hilton. 2010. 3D modelling of static environments using multiple spherical stereo. In Proceedings of the European Conference on Computer Vision, Kiriakos N. Kutulakos (Ed.). 169\u2013183."},{"key":"e_1_3_2_91_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-013-0616-1"},{"key":"e_1_3_2_92_2","first-page":"1","article-title":"Planar urban scene reconstruction from spherical images using facade alignment","author":"Kim H.","year":"2013","unstructured":"H. Kim and A. Hilton. 2013. Planar urban scene reconstruction from spherical images using facade alignment. IEEE IVMSP (2013), 1\u20134.","journal-title":"IEEE IVMSP"},{"key":"e_1_3_2_93_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2015.04.001"},{"key":"e_1_3_2_94_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2017.00076"},{"key":"e_1_3_2_95_2","first-page":"476","volume-title":"Proceedings of the European Conference on Computer Vision.","author":"Ko\u0161eck\u00e1 J.","year":"2002","unstructured":"J. Ko\u0161eck\u00e1 and W. Zhang. 2002. Video compass. In Proceedings of the European Conference on Computer Vision.476\u2013490."},{"issue":"67","key":"e_1_3_2_96_2","article-title":"Spherical light fields","author":"Krolla B.","year":"2014","unstructured":"B. Krolla, M. Diebold, B. Goldl\u00fccke, and D. Stricker. 2014. Spherical light fields. BMVC (2014), 67.1\u201367.12.","journal-title":"BMVC"},{"key":"e_1_3_2_97_2","doi-asserted-by":"publisher","DOI":"10.5815\/ijem.2016.04.05"},{"key":"e_1_3_2_98_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2011.5980103"},{"key":"e_1_3_2_99_2","first-page":"405","article-title":"Real-time panoramic depth maps from omni-directional stereo images for 6 DoF videos in virtual reality","author":"Lai P. K.","year":"2019","unstructured":"P. K. Lai, S. Xie, J. Lang, and R. Laqaruere. 2019. Real-time panoramic depth maps from omni-directional stereo images for 6 DoF videos in virtual reality. IEEE Conference on Virtual Reality and 3D User Interfaces (2019), 405\u2013412.","journal-title":"IEEE Conference on Virtual Reality and 3D User Interfaces"},{"key":"e_1_3_2_100_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00940"},{"key":"e_1_3_2_101_2","first-page":"1","article-title":"SpherePHD: Applying CNNs on 360 \\( ^\\circ \\)  images with non-euclidean spherical PolyHeDron representation","author":"Lee Y.","year":"2020","unstructured":"Y. Lee, J. Jeong, J. Yun, W. Cho, and K.-J. Yoon. 2020. SpherePHD: Applying CNNs on 360 \\( ^\\circ \\) images with non-euclidean spherical PolyHeDron representation. IEEE Transactions on Pattern Analysis and Machine Intelligence (2020), 1\u20131.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_102_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2888856"},{"key":"e_1_3_2_103_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126542"},{"key":"e_1_3_2_104_2","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2006.270"},{"key":"e_1_3_2_105_2","doi-asserted-by":"publisher","DOI":"10.1145\/237170.237199"},{"key":"e_1_3_2_106_2","first-page":"9605","volume-title":"Proceedings of the Advances in Neural Information Processing Systems.","author":"Liu R.","year":"2018","unstructured":"R. Liu, J. Lehman, P. Molino, F. P. Such, E. Frank, A. Sergeev, and J. Yosinski. 2018. An intriguing failing of convolutional neural networks and the CoordConv solution. In Proceedings of the Advances in Neural Information Processing Systems.9605\u20139616."},{"key":"e_1_3_2_107_2","doi-asserted-by":"publisher","DOI":"10.1023\/b:visi.0000029664.99615.94"},{"key":"e_1_3_2_108_2","first-page":"674","volume-title":"Proceedings of the IJCAI","author":"Lucas B. D.","year":"1981","unstructured":"B. D. Lucas and T. Kanade. 1981. An iterative image registration technique with an application to stereo vision. In Proceedings of the IJCAI. 674\u2013679."},{"key":"e_1_3_2_109_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2018.2794071"},{"key":"e_1_3_2_110_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.113"},{"key":"e_1_3_2_111_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206535"},{"key":"e_1_3_2_112_2","first-page":"307","volume-title":"Proceedings of the Workshop of SIMPAR.","author":"Mochizuki Y.","year":"2008","unstructured":"Y. Mochizuki and A. Imiya. 2008. Featureless visual navigation using optical flow of omnidirectional image sequence. In Proceedings of the Workshop of SIMPAR.307\u2013318."},{"key":"e_1_3_2_113_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-23678-5_41"},{"key":"e_1_3_2_114_2","doi-asserted-by":"publisher","DOI":"10.1137\/080732730"},{"key":"e_1_3_2_115_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.1997.609369"},{"key":"e_1_3_2_116_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2004.17"},{"key":"e_1_3_2_117_2","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275031"},{"key":"e_1_3_2_118_2","doi-asserted-by":"publisher","DOI":"10.1017\/S096249291700006X"},{"key":"e_1_3_2_119_2","first-page":"17","volume-title":"Proceedings of the VAST","author":"Pagani A.","year":"2011","unstructured":"A. Pagani, C. Gava, Y. Cui, B. Krolla, J.-M. Hengen, and D. Stricker. 2011. Dense 3D point cloud generation from multiple high-resolution spherical images. In Proceedings of the VAST. 17\u201324."},{"key":"e_1_3_2_120_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2011.6130266"},{"key":"e_1_3_2_121_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAS.2016.7832307"},{"key":"e_1_3_2_122_2","first-page":"887","volume-title":"IEEE\/SICE International Symposium on System Integration","author":"Pathak S.","year":"2017","unstructured":"S. Pathak, A. Moro, H. Fujii, A. Yamashita, and H. Asama. 2017. Virtual reality with motion parallax by dense optical flow-based depth generation from two spherical images. In IEEE\/SICE International Symposium on System Integration. 887\u2013892."},{"key":"e_1_3_2_123_2","doi-asserted-by":"publisher","DOI":"10.1109\/IST.2016.7738212"},{"key":"e_1_3_2_124_2","doi-asserted-by":"publisher","DOI":"10.9746\/jcmsi.10.476"},{"key":"e_1_3_2_125_2","doi-asserted-by":"publisher","DOI":"10.1155\/2017\/3497650"},{"key":"e_1_3_2_126_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58598-3_26"},{"key":"e_1_3_2_127_2","first-page":"45","volume-title":"Proceedings of the PG","author":"Pintore G.","year":"2018","unstructured":"G. Pintore, F. Ganovelli, R. Pintus, R. Scopigno, and E. Gobbetti. 2018. Recovering 3D indoor floor plans by exploiting low-cost spherical photography. In Proceedings of the PG(Short Papers and Posters). 45\u201348."},{"key":"e_1_3_2_128_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2016.7477631"},{"key":"e_1_3_2_129_2","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.14021"},{"key":"e_1_3_2_130_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cag.2018.09.013"},{"key":"e_1_3_2_131_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2011.05.002"},{"key":"e_1_3_2_132_2","doi-asserted-by":"publisher","DOI":"10.5555\/2480985"},{"key":"e_1_3_2_133_2","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Ranjan A.","year":"2016","unstructured":"A. Ranjan and M. J. Black. 2016. Optical flow estimation using a spatial pyramid network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_134_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.91"},{"key":"e_1_3_2_135_2","doi-asserted-by":"publisher","DOI":"10.1007\/11744023_34"},{"key":"e_1_3_2_136_2","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2020.2967657"},{"key":"e_1_3_2_137_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126544"},{"key":"e_1_3_2_138_2","first-page":"217","volume-title":"Proceedings of the VR 2005. Virtual Reality.","author":"Li S.","year":"2005","unstructured":"S. Li and K. Fukumori. 2005. Spherical stereo for the construction of immersive VR environment. In Proceedings of the VR 2005. Virtual Reality.217\u2013222."},{"key":"e_1_3_2_139_2","doi-asserted-by":"publisher","DOI":"10.1145\/3177853"},{"key":"e_1_3_2_140_2","first-page":"1161","volume-title":"Proceedings of the Advances in Neural Information Processing Systems.","author":"Saxena A.","year":"2006","unstructured":"A. Saxena, S. H. Chung, and A. Y. Ng. 2006. Learning depth from single monocular images. In Proceedings of the Advances in Neural Information Processing Systems.1161\u20131168."},{"issue":"1","key":"e_1_3_2_141_2","first-page":"131","article-title":"A taxonomy and evaluation of dense two-frame stereo correspondence algorithms","volume":"47","author":"Scharstein D.","year":"2001","unstructured":"D. Scharstein, R. Szeliski, and R. Zabih. 2001. A taxonomy and evaluation of dense two-frame stereo correspondence algorithms. SMBV 47, 1 (2001), 131\u2013140.","journal-title":"SMBV"},{"key":"e_1_3_2_142_2","volume-title":"Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems.","author":"Sch\u00f6nbein M.","year":"2014","unstructured":"M. Sch\u00f6nbein and A. Geiger. 2014. Omnidirectional 3D reconstruction in augmented Manhattan worlds. In Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems."},{"key":"e_1_3_2_143_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.19"},{"key":"e_1_3_2_144_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.comgeo.2014.02.005"},{"key":"e_1_3_2_145_2","unstructured":"V. Sitzmann S. Rezchikov W. T. Freeman J. B. Tenenbaum and F. Durand. 2021. Light field networks: Neural scene representations with single-evaluation rendering. (2021). arxiv:2106.02634. Retrieved from https:\/\/arxiv.org\/abs\/2106.02634."},{"key":"e_1_3_2_146_2","article-title":"Semantic scene completion from a single depth image","author":"Song S.","year":"2017","unstructured":"S. Song, F. Yu, A. Zeng, A. X. Chang, M. Savva, and T. Funkhouser. 2017. Semantic scene completion from a single depth image. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2017).","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"e_1_3_2_147_2","first-page":"529","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","author":"Su Y.-C.","year":"2017","unstructured":"Y.-C. Su and K. Grauman. 2017. Learning spherical convolution for fast features from 360 \\( \\backslash \\) textdegree imagery. In Proceedings of the Advances in Neural Information Processing Systems. 529\u2013539."},{"key":"e_1_3_2_148_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350539"},{"key":"e_1_3_2_149_2","doi-asserted-by":"crossref","unstructured":"C. Sun C.-W. Hsiao M. Sun and H.-T. Chen. 2019. HorizonNet: Learning room layout with 1D representation and pano stretch data augmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 1047\u20131056.","DOI":"10.1109\/CVPR.2019.00114"},{"key":"e_1_3_2_150_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00260"},{"key":"e_1_3_2_151_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00931"},{"key":"e_1_3_2_152_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2017.2720693"},{"key":"e_1_3_2_153_2","doi-asserted-by":"publisher","DOI":"10.1145\/1882261.1866194"},{"key":"e_1_3_2_154_2","first-page":"732","article-title":"Distortion-aware convolutional filters for dense prediction in panoramic images","author":"Tateno K.","year":"2018","unstructured":"K. Tateno, N. Navab, and F. Tombari. 2018. Distortion-aware convolutional filters for dense prediction in panoramic images. Proceedings of the European Conference on Computer Vision.732\u2013750.","journal-title":"Proceedings of the European Conference on Computer Vision."},{"key":"e_1_3_2_155_2","unstructured":"L. Tchapmi and D. Huber. 2019. The SUMO challenge. The 20 9 (2019) 667\u2013676."},{"key":"e_1_3_2_156_2","doi-asserted-by":"publisher","DOI":"10.1109\/VCIP.2017.8305085"},{"key":"e_1_3_2_157_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2009.5457551"},{"key":"e_1_3_2_158_2","first-page":"1","volume-title":"Proceedings of the 13th European Signal Processing Conference.","author":"Tosic I.","year":"2005","unstructured":"I. Tosic, I. Bogdanova, P. Frossard, and P. Vandergheynst. 2005. Multiresolution motion estimation for omnidirectional images. In Proceedings of the 13th European Signal Processing Conference.1\u20134."},{"key":"e_1_3_2_159_2","doi-asserted-by":"crossref","unstructured":"F.-E. Wang H.-N. Hu H.-T. Cheng J.-T. Lin S.-T. Yang M.-L. Shih H.-K. Chu and M. Sun. 2018. Self-supervised learning of depth and camera motion from 360 \\( ^\\circ \\)  videos. Asian Conference on Computer Vision 11364 (2018) 53\u201368.","DOI":"10.1007\/978-3-030-20873-8_4"},{"key":"e_1_3_2_160_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00054"},{"key":"e_1_3_2_161_2","doi-asserted-by":"crossref","unstructured":"F.-E. Wang Y.-H. Yeh M. Sun W.-C. Chiu and Y.-H. Tsai. 2021. LED2-Net: Monocular 360deg layout estimation via differentiable depth rendering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 12956\u201312965.","DOI":"10.1109\/CVPR46437.2021.01276"},{"key":"e_1_3_2_162_2","article-title":"360SD-Net: 360\u00b0 stereo depth estimation with learnable cost volume","author":"Wang N.-H.","year":"2020","unstructured":"N.-H. Wang, B. Solarte, Y.-H. Tsai, W.-C. Chiu, and M. Sun. 2020. 360SD-Net: 360\u00b0 stereo depth estimation with learnable cost volume. IEEE International Conference on Robotics and Automation (2020), 582\u2013588.","journal-title":"IEEE International Conference on Robotics and Automation"},{"key":"e_1_3_2_163_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8463173"},{"key":"e_1_3_2_164_2","unstructured":"K. Wegner O. Stankiewicz and M. Doma\u0144ski. 2015. Depth based view blending in view synthesis reference software (VSRS). ISO\/IEC JTC1\/SC29\/WG11 MPEG2015 M37232 Geneva Switzerland ."},{"key":"e_1_3_2_165_2","first-page":"2945","article-title":"Depth estimation from stereoscopic 360-degree video","author":"Wegner K.","year":"2018","unstructured":"K. Wegner, O. Stankiewicz, T. Grajek, and M. Domanski. 2018. Depth estimation from stereoscopic 360-degree video. In Proceedings of the 25th IEEE International Conference on Image Processing.2945\u20132948.","journal-title":"Proceedings of the 25th IEEE International Conference on Image Processing."},{"key":"e_1_3_2_166_2","first-page":"1385","article-title":"DeepFlow: Large displacement optical flow with deep matching","author":"Weinzaepfel P.","year":"2013","unstructured":"P. Weinzaepfel, J. Revaud, Z. Harchaoui, and C. Schmid. 2013. DeepFlow: Large displacement optical flow with deep matching. In Proceedings of the IEEE International Conference on Computer Vision.1385\u20131392.","journal-title":"Proceedings of the IEEE International Conference on Computer Vision."},{"key":"e_1_3_2_167_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2898948"},{"key":"e_1_3_2_168_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00908"},{"key":"e_1_3_2_169_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8793823"},{"key":"e_1_3_2_170_2","article-title":"End-to-end learning for omnidirectional stereo matching with uncertainty prior","author":"Won C.","year":"2020","unstructured":"C. Won, J. Ryu, and J. Lim. 2020. End-to-end learning for omnidirectional stereo matching with uncertainty prior. IEEE Transactions on Pattern Analysis and Machine Intelligence (2020).","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_171_2","doi-asserted-by":"crossref","unstructured":"C. Won H. Seok Z. Cui M. Pollefeys and J. Lim. 2020. OmniSLAM: Omnidirectional localization and dense mapping for wide-baseline multi-camera systems. In Proceedings of the IEEE International Conference on Robotics and Automation. 559\u2013566.","DOI":"10.1109\/ICRA40945.2020.9196695"},{"key":"e_1_3_2_172_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2019.00079"},{"key":"e_1_3_2_173_2","doi-asserted-by":"publisher","DOI":"10.1109\/ROBIO.2016.7866353"},{"key":"e_1_3_2_174_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2017.46"},{"key":"e_1_3_2_175_2","doi-asserted-by":"publisher","DOI":"10.1109\/CADGraphics.2013.77"},{"key":"e_1_3_2_176_2","doi-asserted-by":"publisher","DOI":"10.1145\/2670473.2670485"},{"key":"e_1_3_2_177_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.585"},{"key":"e_1_3_2_178_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00348"},{"key":"e_1_3_2_179_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00413"},{"key":"e_1_3_2_180_2","doi-asserted-by":"crossref","unstructured":"C. Zach T. Pock and H. Bischof. 2007. A duality based approach for realtime TV-L1 optical flow. In Proceedings of the Pattern Recognition. 214\u2013223.","DOI":"10.1007\/978-3-540-74936-3_22"},{"key":"e_1_3_2_181_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.107861"},{"key":"e_1_3_2_182_2","first-page":"7324","volume-title":"Proceedings of the International Conference on Machine Learning.","author":"Zhang R.","year":"2019","unstructured":"R. Zhang. 2019. Making convolutional networks shift-invariant again. In Proceedings of the International Conference on Machine Learning.7324\u20137334."},{"key":"e_1_3_2_183_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10599-4_43"},{"key":"e_1_3_2_184_2","first-page":"801","article-title":"Benefit of large field-of-view cameras for visual odometry","author":"Zhang Z.","year":"2016","unstructured":"Z. Zhang, H. Rebecq, C. Forster, and D. Scaramuzza. 2016. Benefit of large field-of-view cameras for visual odometry. In Proceedings of the IEEE International Conference on Robotics and Automation.801\u2013808.","journal-title":"Proceedings of the IEEE International Conference on Robotics and Automation."},{"key":"e_1_3_2_185_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-20351-1_67"},{"key":"e_1_3_2_186_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-014-0787-4"},{"key":"e_1_3_2_187_2","doi-asserted-by":"crossref","unstructured":"J. Zheng J. Zhang J. Li R. Tang S. Gao and Z. Zhou. 2020. Structured3D: A large photo-realistic dataset for structured 3D modeling. In Proceedings of the European Conference on Computer Vision. 519\u2013535.","DOI":"10.1007\/978-3-030-58545-7_30"},{"key":"e_1_3_2_188_2","doi-asserted-by":"crossref","unstructured":"N. Zioulis F. Alvarez D. Zarpalas and P. Daras. 2021. Single-shot cuboids: Geodesics-based end-to-end manhattan aligned layout estimation from spherical panoramas. Image and Vision Computing 110 (2021) 104160 pages.","DOI":"10.1016\/j.imavis.2021.104160"},{"key":"e_1_3_2_189_2","first-page":"690","article-title":"Spherical view synthesis for self-supervised 360 \\( ^\\circ \\)  depth estimation","author":"Zioulis N.","year":"2019","unstructured":"N. Zioulis, A. Karakottas, D. Zarpalas, F. Alvarez, and P. Daras. 2019. Spherical view synthesis for self-supervised 360 \\( ^\\circ \\) depth estimation. InProceedings of the International Conference on 3D Vision.690\u2013699.","journal-title":"Proceedings of the International Conference on 3D Vision."},{"key":"e_1_3_2_190_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01231-1_28"},{"key":"e_1_3_2_191_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00219"},{"key":"e_1_3_2_192_2","doi-asserted-by":"crossref","unstructured":"C. Zou J.-W. Su C.-H. Peng A. Colburn Q. Shan P. Wonka H.-K. Chu and D. Hoiem. 2021. Manhattan room layout reconstruction from a single 360 \\( ^\\circ \\)  image: A comparative study of state-of-the-art methods. International Journal of Computer Vision (2021) 1\u201322.","DOI":"10.1007\/s11263-020-01426-8"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3519021","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3519021","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:12:20Z","timestamp":1750191140000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3519021"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,11,21]]},"references-count":191,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,4,30]]}},"alternative-id":["10.1145\/3519021"],"URL":"https:\/\/doi.org\/10.1145\/3519021","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,11,21]]},"assertion":[{"value":"2021-04-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-02-14","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-11-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}