{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T12:58:19Z","timestamp":1772283499057,"version":"3.50.1"},"reference-count":104,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,2,6]],"date-time":"2026-02-06T00:00:00Z","timestamp":1770336000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,2,9]],"date-time":"2026-02-09T00:00:00Z","timestamp":1770595200000},"content-version":"vor","delay-in-days":3,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003696","name":"Electronics and Telecommunications Research Institute","doi-asserted-by":"publisher","award":["24ZC1200"],"award-info":[{"award-number":["24ZC1200"]}],"id":[{"id":"10.13039\/501100003696","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Virtual Reality"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    We introduce a new approach for constructing immersive virtual spaces by generating comprehensive 3D voxelised models that encompass both geometric and semantic scene representations from a single 360\n                    <jats:inline-formula>\n                      <jats:tex-math>$${}^{\\circ}$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    RGB-D input. The proposed approach utilises a deep convolutional neural network for semantic scene completion (SSC), allowing the estimation of complete semantics and geometries of the scene. We design MDBNet a dual head model that simultaneously processes RGB and depth data using a perspective camera. Depth information is encoded using a flipped transcribed signed distance function (F-TSDF), capturing essential geometric shape characteristics. We extend the inference capabilities of MDBNet on RGB-D input of the perspective camera to accommodate 360\n                    <jats:inline-formula>\n                      <jats:tex-math>$${}^{\\circ}$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    RGB-D by proposing MDBNet360. We employ RGB spherical-to-cubic projection and 3D rotation for depth point clouds, allowing for virtual reality (VR) space design with comprehensive spatial coverage. To our knowledge, this is the first work to extend a pre-trained SSC model, originally using perspective camera RGB-D input, to infer a 3D model from 360\n                    <jats:inline-formula>\n                      <jats:tex-math>$${}^{\\circ }$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    RGB-D input. To assess acoustic properties, we measure parameters such as early decay time (EDT) and reverberation time (RT60) using the exponential sine sweep method (ESS). We used Unity with the Steam Audio plug-in for conducting simulations in virtual space. The proposed framework demonstrates better virtual space reconstruction and immersive sound generation, advancing semantically rich and spatially accurate virtual environments compared to the state-of-the-art (SOTA). Code and rendered sounds are available on GitHub:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/MonaIA1\/Repo360\" ext-link-type=\"uri\">https:\/\/github.com\/MonaIA1\/Repo360<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s10055-026-01312-7","type":"journal-article","created":{"date-parts":[[2026,2,6]],"date-time":"2026-02-06T10:55:48Z","timestamp":1770375348000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["3D audio-visual indoor scene reconstruction and semantics completion for virtual reality from a single $$360^{\\circ }$$ RGB-D image"],"prefix":"10.1007","volume":"30","author":[{"given":"Mona","family":"Alawadh","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Atiyeh","family":"Alinaghi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mahesan","family":"Niranjan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hansung","family":"Kim","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,2,6]]},"reference":[{"key":"1312_CR1","unstructured":"ADE20K dataset (2023). https:\/\/tinyurl.com\/ADE20K. Accessed 17 Jan 2023"},{"key":"1312_CR2","doi-asserted-by":"crossref","unstructured":"Alawadh M, Niranjan M, Kim H (2024) 3d semantic scene completion from a depth map with unsupervised learning for semantics prioritisation. In: 2024 IEEE International Conference on Image Processing (ICIP), IEEE, pp 3348\u20133354","DOI":"10.1109\/ICIP51287.2024.10647579"},{"key":"1312_CR3","unstructured":"Anil \u00c7 (2024) Modern workflows for procedural audio at the intersection of gaming and music performance in virtual reality. In: Audio Engineering Society Conference: AES 2024 International Audio for Games Conference. Audio Engineering Society"},{"key":"1312_CR4","unstructured":"Armeni I, Sax S, Zamir AR, Savarese S (2017) Joint 2d-3d-semantic data for indoor scene understanding. arXiv preprint arXiv:1702.01105"},{"issue":"12","key":"1312_CR5","doi-asserted-by":"publisher","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","volume":"39","author":"V Badrinarayanan","year":"2017","unstructured":"Badrinarayanan V, Kendall A, Cipolla R (2017) Segnet: a deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans Pattern Anal Mach Intell 39(12):2481\u20132495","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"3\u2013Supplement","key":"1312_CR6","doi-asserted-by":"publisher","first-page":"282","DOI":"10.1121\/10.0027512","volume":"155","author":"M-V Baran","year":"2024","unstructured":"Baran M-V, King R, Woszczyk W (2024) A general overview of methods for generating room impulse responses. J Acoust Soc Am 155(3\u2013Supplement):282\u2013282","journal-title":"J Acoust Soc Am"},{"issue":"4","key":"1312_CR7","first-page":"320","volume":"81","author":"M Barron","year":"1995","unstructured":"Barron M (1995) Interpretation of early decay times in concert auditoria. Acta Acust Acust 81(4):320\u2013331","journal-title":"Acta Acust Acust"},{"key":"1312_CR8","doi-asserted-by":"publisher","first-page":"873","DOI":"10.1007\/978-3-031-23161-2_169","volume-title":"Encyclopedia of computer graphics and games","author":"MI Berkman","year":"2024","unstructured":"Berkman MI (2024) History of virtual reality. In: Lee N (ed) Encyclopedia of computer graphics and games. Springer, Cham, pp 873\u2013881"},{"issue":"10","key":"1312_CR9","doi-asserted-by":"publisher","first-page":"713","DOI":"10.1016\/j.apacoust.2011.04.004","volume":"72","author":"JS Bradley","year":"2011","unstructured":"Bradley JS (2011) Review of objective room acoustics measures and future needs. Appl Acoust 72(10):713\u2013720","journal-title":"Appl Acoust"},{"key":"1312_CR10","doi-asserted-by":"crossref","unstructured":"Cai Y, Chen X, Zhang C, Lin K-Y, Wang X, Li H (2021) Semantic scene completion via integrating instances and scene in-the-loop. In: CVPR, pp 324\u2013333","DOI":"10.1109\/CVPR46437.2021.00039"},{"key":"1312_CR11","doi-asserted-by":"crossref","unstructured":"Cao A-Q, Charette R (2022) Monoscene: monocular 3d semantic scene completion. In: CVPR, pp 3991\u20134001","DOI":"10.1109\/CVPR52688.2022.00396"},{"key":"1312_CR12","unstructured":"Cao K, Wei C, Gaidon A, Arechiga N, Ma T (2019) Learning imbalanced datasets with label-distribution-aware margin loss. Adv Neural Inf Process Syst 32"},{"issue":"17","key":"1312_CR13","doi-asserted-by":"publisher","first-page":"1658","DOI":"10.1080\/10447318.2020.1778351","volume":"36","author":"E Chang","year":"2020","unstructured":"Chang E, Kim HT, Yoo B (2020) Virtual reality sickness: a review of causes and measurements. Int J Human-Comput Interact 36(17):1658\u20131682","journal-title":"Int J Human-Comput Interact"},{"key":"1312_CR14","doi-asserted-by":"crossref","unstructured":"Chang A, Dai A, Funkhouser T, Halber M, Niessner M, Savva M, Song S, Zeng A, Zhang Y (2017) Matterport3d: learning from rgb-d data in indoor environments. arXiv preprint arXiv:1709.06158","DOI":"10.1109\/3DV.2017.00081"},{"key":"1312_CR15","doi-asserted-by":"crossref","unstructured":"Chen L-C, Zhu Y, Papandreou G, Schroff F, Adam H (2018) Encoder-decoder with atrous separable convolution for semantic image segmentation. In: ECCV, pp 801\u2013818","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"1312_CR16","doi-asserted-by":"crossref","unstructured":"Chen X, Lin K-Y, Qian C, Zeng , Li H (2020) 3d sketch-aware semantic scene completion via semi-supervised structure prior. In: CVPR, pp 4193\u20134202","DOI":"10.1109\/CVPR42600.2020.00425"},{"key":"1312_CR17","doi-asserted-by":"crossref","unstructured":"Chen M, Su K, Shlizerman E (2023) Be everywhere-hear everything (bee): audio scene reconstruction by sparse audio-visual samples. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp 7853\u20137862","DOI":"10.1109\/ICCV51070.2023.00722"},{"key":"1312_CR18","doi-asserted-by":"publisher","first-page":"247","DOI":"10.35784\/jcsi.2698","volume":"20","author":"A Ciekanowska","year":"2021","unstructured":"Ciekanowska A, Kiszczak-Gli\u0144ski A, Dziedzic K (2021) Vr space such as. J Comput Sci Inst 20:247\u2013253","journal-title":"J Comput Sci Inst"},{"issue":"4","key":"1312_CR19","doi-asserted-by":"publisher","first-page":"77","DOI":"10.3390\/technologies8040077","volume":"8","author":"S Doolani","year":"2020","unstructured":"Doolani S, Wessels C, Kanal V, Sevastopoulos C, Jaiswal A, Nambiappan H, Makedon F (2020) A review of extended reality (xr) technologies for manufacturing training. Technologies 8(4):77","journal-title":"Technologies"},{"key":"1312_CR20","doi-asserted-by":"crossref","unstructured":"Dourado A, De Campos TE, Kim H, Hilton A (2021) Edgenet: semantic scene completion from a single rgb-d image. In: ICPR, pp 503\u2013510","DOI":"10.1109\/ICPR48806.2021.9413252"},{"key":"1312_CR21","doi-asserted-by":"crossref","unstructured":"Dourado A, Guth F, Campos T (2022) Data augmented 3d semantic scene completion with 2d segmentation priors. In: IEEE Winter Conference on Applications of Computer Vision (WACV), pp 3781\u20133790","DOI":"10.1109\/WACV51458.2022.00076"},{"key":"1312_CR22","volume-title":"Springer handbook of acoustics","author":"F Dunn","year":"2015","unstructured":"Dunn F, Hartmann W, Campbell D, Fletcher NH (2015) Springer handbook of acoustics. Springer, New York"},{"key":"1312_CR23","unstructured":"Farina A (2000) Simultaneous measurement of impulse response and distortion with a swept-sine technique. In: Audio Engineering Society Convention 108, Audio Engineering Society"},{"key":"1312_CR24","unstructured":"Farina A (2007) Advancements in impulse response measurements by sine sweeps. In: Audio Engineering Society Convention 122, Audio Engineering Society"},{"key":"1312_CR25","doi-asserted-by":"crossref","unstructured":"Firman M, Mac Aodha O, Julier S, Brostow GJ (2016) Structured prediction of unobserved voxels from a single depth image. In: CVPR, pp 5431\u20135440","DOI":"10.1109\/CVPR.2016.586"},{"key":"1312_CR26","doi-asserted-by":"crossref","unstructured":"Garbade M, Chen Y-T, Sawatzky J, Gall J (2019) Two stream 3d semantic scene completion. In: CVPRW, pp 0\u20130","DOI":"10.1109\/CVPRW.2019.00055"},{"key":"1312_CR27","doi-asserted-by":"crossref","unstructured":"Han H, Liang Y, Zhou Y, Wang W, J.\u00a0Rojas-Mu\u00f1oz E, Li X (2024) Aurora: automated unleash of 3d room outlines for vr applications. In: Proceedings of the 19th ACM SIGGRAPH International Conference on Virtual-Reality Continuum and Its Applications in Industry, pp 1\u20138","DOI":"10.1145\/3703619.3706036"},{"key":"1312_CR28","doi-asserted-by":"crossref","unstructured":"Heng Y, Dasmahapatra S, Kim H (2023) Dbat: dynamic backward attention transformer for material segmentation with cross-resolution patches. arXiv preprint arXiv:2305.03919","DOI":"10.2139\/ssrn.4860829"},{"key":"1312_CR29","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: CVPR, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"1312_CR30","unstructured":"International Organization for Standardization: ISO 3382-1:2009: Acoustics \u2013 measurement of room acoustic parameters \u2013 part 1: performance spaces. https:\/\/www.iso.org\/standard\/40979.html"},{"key":"1312_CR31","unstructured":"IoSR: IoSR Matlab Toolbox. https:\/\/github.com\/IoSR-Surrey\/MatlabToolbox\/tree\/master. Accessed: 2 Sept 2024"},{"issue":"3","key":"1312_CR32","doi-asserted-by":"publisher","first-page":"14","DOI":"10.12948\/issn14531305\/22.3.2018.02","volume":"22","author":"C Isar","year":"2018","unstructured":"Isar C (2018) A glance into virtual reality development using unity. Inf Economica 22(3):14\u201322","journal-title":"Inf Economica"},{"issue":"4","key":"1312_CR33","doi-asserted-by":"publisher","first-page":"396","DOI":"10.9734\/BJAST\/2015\/14975","volume":"7","author":"A Joshi","year":"2015","unstructured":"Joshi A, Kale S, Chandel S, Pal DK (2015) Likert scale: explored and explained. British J Appl Sci Technol 7(4):396","journal-title":"British J Appl Sci Technol"},{"key":"1312_CR34","doi-asserted-by":"publisher","first-page":"104","DOI":"10.1016\/j.cviu.2015.04.001","volume":"139","author":"H Kim","year":"2015","unstructured":"Kim H, Hilton A (2015) Block world reconstruction from spherical stereo image pairs. Comput Vis Image Underst 139:104\u2013121","journal-title":"Comput Vis Image Underst"},{"key":"1312_CR35","doi-asserted-by":"publisher","first-page":"293","DOI":"10.1007\/978-3-030-41816-8_13","volume-title":"Real vr-immersive digital reality: how to import the real world into head-mounted immersive displays","author":"H Kim","year":"2020","unstructured":"Kim H, Remaggi L, Jackson PJB, Hilton A (2020) Immersive virtual reality audio rendering adapted to the listener and the room. In: Magnor M, Sorkine-Hornung A (eds) Real vr-immersive digital reality: how to import the real world into head-mounted immersive displays. Springer, Cham, pp 293\u2013318"},{"issue":"3","key":"1312_CR36","doi-asserted-by":"publisher","first-page":"823","DOI":"10.1007\/s10055-021-00594-3","volume":"26","author":"H Kim","year":"2022","unstructured":"Kim H, Remaggi L, Dourado A, Campos TD, Jackson PJ, Hilton A (2022) Immersive audio-visual scene reproduction using semantic scene reconstruction from 360 cameras. Virtual Real 26(3):823\u2013838","journal-title":"Virtual Real"},{"key":"1312_CR37","doi-asserted-by":"crossref","unstructured":"Kim H, Remaggi L, Jackson PJ, Hilton A (2019) Immersive spatial audio reproduction for vr\/ar using room acoustic modelling from 360 images. In: 2019 IEEE Conference on Virtual Reality and 3D User Interfaces (VR), pp 120\u2013126","DOI":"10.1109\/VR.2019.8798247"},{"issue":"3","key":"1312_CR38","doi-asserted-by":"publisher","first-page":"226","DOI":"10.1109\/34.667881","volume":"20","author":"J Kittler","year":"1998","unstructured":"Kittler J, Hatef M, Duin RP, Matas J (1998) On combining classifiers. IEEE Trans Pattern Anal Mach Intell 20(3):226\u2013239","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"1312_CR39","unstructured":"Kon H, Koike H (2018) Deep neural networks for cross-modal estimations of acoustic reverberation characteristics from two-dimensional images. In: Audio Engineering Society Convention 144"},{"key":"1312_CR40","first-page":"57050","volume":"37","author":"S Lee","year":"2024","unstructured":"Lee S, Chung J, Huh J, Lee KM (2024) ODGS: 3d scene reconstruction from omnidirectional images with 3d gaussian splattings. Adv Neural Inf Process Syst 37:57050\u201357075","journal-title":"Adv Neural Inf Process Syst"},{"key":"1312_CR41","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1155\/2007\/70540","volume":"2007","author":"T Lentz","year":"2007","unstructured":"Lentz T, Schr\u00f6der D, Vorl\u00e4nder M, Assenmacher I (2007) Virtual reality system with integrated sound field simulation and reproduction. EURASIP J Adv Signal Process 2007:1\u201319","journal-title":"EURASIP J Adv Signal Process"},{"issue":"1","key":"1312_CR42","doi-asserted-by":"publisher","first-page":"219","DOI":"10.1109\/LRA.2019.2953639","volume":"5","author":"J Li","year":"2019","unstructured":"Li J, Liu Y, Yuan X, Zhao C, Siegwart R, Reid I, Cadena C (2019) Depth based semantic scene completion with position importance aware loss. IEEE Robot Automat Lett 5(1):219\u2013226","journal-title":"IEEE Robot Automat Lett"},{"key":"1312_CR43","unstructured":"Liang S, Huang C, Tian Y, Kumar A, Xu C (2023) Neural acoustic context field: rendering realistic room impulse response with neural fields. arXiv preprint arXiv:2309.15977"},{"key":"1312_CR44","doi-asserted-by":"crossref","unstructured":"Li J, Ding L, Huang R (2021) Imenet: joint 3d semantic scene completion and 2d semantic segmentation through iterative mutual enhancement. In: IJCAI","DOI":"10.24963\/ijcai.2021\/110"},{"key":"1312_CR45","doi-asserted-by":"crossref","unstructured":"Li J, Han K, Wang P, Liu Y, Yuan X (2020) Anisotropic convolutional networks for 3d semantic scene completion. In: CVPR, pp 3351\u20133359","DOI":"10.1109\/CVPR42600.2020.00341"},{"key":"1312_CR46","doi-asserted-by":"crossref","unstructured":"Li J, Liu Y, Gong D, Shi Q, Yuan X, Zhao C, Reid I (2019) Rgbd based dimensional decomposition residual network for 3d semantic scene completion. In: CVPR, pp 7693\u20137702","DOI":"10.1109\/CVPR.2019.00788"},{"key":"1312_CR47","doi-asserted-by":"crossref","unstructured":"Li M, Meng M, Zhou Z (2022) Repf-net: distortion-aware re-projection fusion network for object detection in panorama image. In: Proceedings of the Asian Conference on Computer Vision, pp 74\u201389","DOI":"10.1007\/978-3-031-26313-2_31"},{"key":"1312_CR48","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Goyal P, Girshick R, He K, Doll\u00e1r P (2017) Focal loss for dense object detection. In: ICCV, pp 2980\u20132988","DOI":"10.1109\/ICCV.2017.324"},{"key":"1312_CR49","first-page":"8294","volume":"25","author":"J Li","year":"2023","unstructured":"Li J, Song Q, Yan X, Chen Y, Huang R (2023) From front to rear: 3d semantic scene completion through planar convolution and attention-based network. IEEE TMM 25:8294\u20138307","journal-title":"IEEE TMM"},{"key":"1312_CR50","unstructured":"Liu S, Hu Y, Zeng Y, Tang Q, Jin B, Han Y, Li X (2018) See and think: disentangling semantic scene completion 31"},{"issue":"3","key":"1312_CR51","doi-asserted-by":"publisher","first-page":"1306","DOI":"10.1007\/s11263-024-02244-y","volume":"133","author":"X Liu","year":"2024","unstructured":"Liu X, Xie H, Zhang S, Yao H, Ji R, Nie L, Tao D (2024) 2d semantic-guided semantic scene completion. Int J Comput Vision 133(3):1306\u20131325","journal-title":"Int J Comput Vision"},{"issue":"3","key":"1312_CR52","doi-asserted-by":"publisher","first-page":"1891","DOI":"10.1007\/s00371-024-03509-w","volume":"41","author":"T Li","year":"2024","unstructured":"Li T, Zhang Z, Wang Y, Cui Y, Li Y, Zhou D, Yin B, Yang X (2024) Self-supervised indoor scene point cloud completion from a single panorama. Visual Comput 41(3):1891\u20131905","journal-title":"Visual Comput"},{"key":"1312_CR53","doi-asserted-by":"crossref","unstructured":"Li S, Zou C, Li Y, Zhao X, Gao Y (2020) Attention-based multi-modal fusion network for semantic scene completion. In: AAAI, pp 11402\u201311409","DOI":"10.1609\/aaai.v34i07.6803"},{"key":"1312_CR54","first-page":"2522","volume":"35","author":"S Majumder","year":"2022","unstructured":"Majumder S, Chen C, Al-Halah Z, Grauman K (2022) Few-shot audio-visual learning of environment acoustics. Adv Neural Inf Process Syst 35:2522\u20132536","journal-title":"Adv Neural Inf Process Syst"},{"issue":"4","key":"1312_CR55","first-page":"304","volume":"4","author":"S Mandal","year":"2013","unstructured":"Mandal S (2013) Brief introduction of virtual reality & its challenges. Int J Sci Eng Res 4(4):304\u2013309","journal-title":"Int J Sci Eng Res"},{"issue":"7","key":"1312_CR56","doi-asserted-by":"publisher","DOI":"10.1016\/j.jksuci.2024.102151","volume":"36","author":"M Meng","year":"2024","unstructured":"Meng M, Zhou Y, Zuo D, Li Z, Zhou Z (2024) Structure recovery from single omnidirectional image with distortion-aware learning. J King Saud Univ-Comput Inf Sci 36(7):102151","journal-title":"J King Saud Univ-Comput Inf Sci"},{"key":"1312_CR57","doi-asserted-by":"crossref","unstructured":"Meng Z, Zhao F, He M (2006) The just noticeable difference of noise length and reverberation perception. In: 2006 International Symposium on Communications and Information Technologies, IEEE, pp 418\u2013421","DOI":"10.1109\/ISCIT.2006.339980"},{"key":"1312_CR58","unstructured":"Mo\u010dnik M (2023) Pyrirtool: a python tool for room impulse response (RIR) processing. https:\/\/github.com\/maj4e\/pyrirtool. Accessed: 2 Sept 2024"},{"issue":"6","key":"1312_CR59","doi-asserted-by":"publisher","first-page":"3947","DOI":"10.1007\/s10462-019-09784-7","volume":"53","author":"R Moradi","year":"2020","unstructured":"Moradi R, Berangi R, Minaei B (2020) A survey of regularization strategies for deep models. Artif Intell Rev 53(6):3947\u20133986","journal-title":"Artif Intell Rev"},{"key":"1312_CR60","unstructured":"NVIDIA: SegFormer B5 Finetuned ADE 640x640. http:\/\/tinyurl.com\/segformerb5. Accessed: 6 Feb 2024"},{"issue":"7","key":"1312_CR61","doi-asserted-by":"publisher","first-page":"6955","DOI":"10.1109\/TITS.2023.3256442","volume":"24","author":"Y Pan","year":"2023","unstructured":"Pan Y, Xie F, Zhao H (2023) Understanding the challenges when 3d semantic segmentation faces class imbalanced and ood data. IEEE Trans Intell Transp Syst 24(7):6955\u20136970","journal-title":"IEEE Trans Intell Transp Syst"},{"key":"1312_CR62","doi-asserted-by":"crossref","unstructured":"Park JJ, Florence PR, Straub J, Newcombe RA, Lovegrove S (2019) Deepsdf: learning continuous signed distance functions for shape representation. 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 165\u2013174","DOI":"10.1109\/CVPR.2019.00025"},{"issue":"2","key":"1312_CR63","doi-asserted-by":"publisher","first-page":"269","DOI":"10.3390\/electronics13020269","volume":"13","author":"N Partarakis","year":"2024","unstructured":"Partarakis N, Zabulis X (2024) A review of immersive technologies, knowledge representation, and ai for human-centered digital experiences. Electronics 13(2):269","journal-title":"Electronics"},{"key":"1312_CR64","doi-asserted-by":"crossref","unstructured":"Pi H, Tian S, Lu M, Liu J, Guo Y, Zhang S (2023) A comprehensive comparison of projections in omnidirectional super-resolution. In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, pp 1\u20135","DOI":"10.1109\/ICASSP49357.2023.10096834"},{"key":"1312_CR65","unstructured":"Politis A, Tervo S, Lokki T, Pulkki V (2018) Parametric multidirectional decomposition of microphone recordings for broadband high-order ambisonic encoding. In: Audio Engineering Society Convention 144, Audio Engineering Society"},{"issue":"16","key":"1312_CR66","doi-asserted-by":"publisher","first-page":"16105","DOI":"10.1007\/s11042-024-19288-4","volume":"84","author":"AG Privitera","year":"2024","unstructured":"Privitera AG, Fontana F, Geronazzo M (2024) The role of audio in immersive storytelling: a systematic review in cultural heritage. Multimedia Tools Appl 84(16):16105\u201316143","journal-title":"Multimedia Tools Appl"},{"key":"1312_CR67","doi-asserted-by":"crossref","unstructured":"Raghuvanshi N, Snyder J, Mehra R, Lin M, Govindaraju N (2010) Precomputed wave simulation for real-time sound propagation of dynamic sources in complex scenes. In: ACM SIGGRAPH 2010 Papers, pp 1\u201311","DOI":"10.1145\/1833349.1778805"},{"key":"1312_CR68","doi-asserted-by":"crossref","unstructured":"Ratnarajah A, Ghosh S, Kumar S, Chiniya P, Manocha D (2024) Av-rir: audio-visual room impulse response estimation. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp 27164\u201327175","DOI":"10.1109\/CVPR52733.2024.02565"},{"key":"1312_CR69","unstructured":"Remaggi L, Jackson P, Coleman P (2015) Estimation of room reflection parameters for a reverberant spatial audio object. In: Audio Engineering Society Convention 138"},{"key":"1312_CR70","doi-asserted-by":"crossref","unstructured":"Ridnik T, Ben-Baruch E, Zamir N, Noy A, Friedman I, Protter M, Zelnik-Manor L (2021) Asymmetric loss for multi-label classification. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp 82\u201391","DOI":"10.1109\/ICCV48922.2021.00015"},{"issue":"3","key":"1312_CR71","doi-asserted-by":"publisher","first-page":"569","DOI":"10.1109\/TPAMI.2009.187","volume":"32","author":"JD Rodriguez","year":"2009","unstructured":"Rodriguez JD, Perez A, Lozano JA (2009) Sensitivity analysis of k-fold cross validation in prediction error estimation. IEEE Trans Pattern Anal Mach Intell 32(3):569\u2013575","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"8","key":"1312_CR72","doi-asserted-by":"publisher","first-page":"1978","DOI":"10.1007\/s11263-021-01504-5","volume":"130","author":"L Roldao","year":"2022","unstructured":"Roldao L, De Charette R, Verroust-Blondet A (2022) 3d semantic scene completion: a survey. IJCV 130(8):1978\u20132005","journal-title":"IJCV"},{"key":"1312_CR73","unstructured":"R\u00f8svik PM (2024) Creating a virtual reality orchestral concert experience with 3d audio. Master\u2019s thesis, The University of Bergen"},{"issue":"4","key":"1312_CR74","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2947508","volume":"13","author":"A Rungta","year":"2016","unstructured":"Rungta A, Rust S, Morales N, Klatzky R, Lin M, Manocha D (2016) Psychoacoustic characterization of propagation effects in virtual environments. ACM Trans Appl Percept (TAP) 13(4):1\u201318","journal-title":"ACM Trans Appl Percept (TAP)"},{"issue":"3","key":"1312_CR75","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky O, Deng J, Su H, Krause J, Satheesh S, Ma S, Huang Z, Karpathy A, Khosla A, Bernstein M et al (2015) Imagenet large scale visual recognition challenge. IJCV 115(3):211\u2013252","journal-title":"IJCV"},{"key":"1312_CR76","unstructured":"Sabir A, Hussain R, Pedro A, Soltani M, Lee D, Park C, Pyeon J-H (2024) Synthetic data generation with unity 3d and unreal engine for construction hazard scenarios: a comparative analysis"},{"issue":"3","key":"1312_CR77","doi-asserted-by":"publisher","first-page":"215","DOI":"10.1121\/1.5092821","volume":"145","author":"L Shtrepi","year":"2019","unstructured":"Shtrepi L (2019) Investigation on the diffusive surface modeling detail in geometrical acoustics based simulations. J Acoust Soc Am 145(3):215\u2013221","journal-title":"J Acoust Soc Am"},{"key":"1312_CR78","doi-asserted-by":"crossref","unstructured":"Silberman N, Hoiem D, Kohli P, Fergus R (2012) Indoor segmentation and support inference from rgbd images. In: ECCV, pp 746\u2013760","DOI":"10.1007\/978-3-642-33715-4_54"},{"key":"1312_CR79","doi-asserted-by":"crossref","unstructured":"Singh N, Mentch J, Ng J, Beveridge M, Drori I (2021) Image2reverb: cross-modal reverb impulse response synthesis. In: ICCV, pp 286\u2013295","DOI":"10.1109\/ICCV48922.2021.00035"},{"key":"1312_CR80","doi-asserted-by":"crossref","unstructured":"Song S, Yu F, Zeng A, Chang AX, Savva M, Funkhouser T (2017) Semantic scene completion from a single depth image. In: CVPR, pp 1746\u20131754","DOI":"10.1109\/CVPR.2017.28"},{"key":"1312_CR81","unstructured":"Stecker GC, Moore TM, Folkerts M, Zotkin D, Duraiswami R (2018) Toward objective measures of auditory co-immersion in virtual and augmented reality. In: Audio Engineering Society Conference: 2018 AES International Conference on Audio for Virtual and Augmented Reality, Audio Engineering Society"},{"issue":"2","key":"1312_CR82","doi-asserted-by":"publisher","first-page":"111","DOI":"10.1111\/j.2517-6161.1974.tb00994.x","volume":"36","author":"M Stone","year":"1974","unstructured":"Stone M (1974) Cross-validatory choice and assessment of statistical predictions. J Roy Stat Soc: Ser B (Methodol) 36(2):111\u2013133","journal-title":"J Roy Stat Soc: Ser B (Methodol)"},{"key":"1312_CR83","doi-asserted-by":"crossref","unstructured":"Tang J, Chen X, Wang J, Zeng G (2022) Not all voxels are equal: semantic scene completion from the point-voxel perspective. In: AAAI, pp 2352\u20132360","DOI":"10.1609\/aaai.v36i2.20134"},{"issue":"11","key":"1312_CR84","doi-asserted-by":"publisher","first-page":"1797","DOI":"10.1109\/TVCG.2012.27","volume":"18","author":"M Taylor","year":"2012","unstructured":"Taylor M, Chandak A, Mo Q, Lauterbach C, Schissler C, Manocha D (2012) Guided multiview ray tracing for fast auralization. IEEE Trans Visual Comput Graph 18(11):1797\u20131810","journal-title":"IEEE Trans Visual Comput Graph"},{"issue":"1","key":"1312_CR85","first-page":"52","volume":"4","author":"R Torres","year":"2004","unstructured":"Torres R, Rycker N, Kleiner M (2004) Edge diffraction and surface scattering in concert halls: physical and perceptual aspects. J Temp Design Architect Environ 4(1):52\u201358","journal-title":"J Temp Design Architect Environ"},{"key":"1312_CR86","doi-asserted-by":"crossref","unstructured":"Van Damme S, Vega MT, De Turck F (2020) Human-centric quality management of immersive multimedia applications. In: 2020 6th IEEE Conference on Network Softwarization (NetSoft), IEEE, pp 57\u201364","DOI":"10.1109\/NetSoft48620.2020.9165335"},{"key":"1312_CR87","unstructured":"Vorl\u00e4nder M (1995) International round robin on room acoustical computer simulations. In: Proceedings of the 15th International Congress on Acoustics (ICA), Trondheim, Norway"},{"key":"1312_CR88","doi-asserted-by":"crossref","unstructured":"Wang Y (2024) Projection methods for 360-degree video. Front Comput Intell Syst","DOI":"10.54097\/e1401963"},{"key":"1312_CR89","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2023.122885","volume":"243","author":"X Wang","year":"2024","unstructured":"Wang X, Feng W, Wan L (2024) Multi-modal fusion architecture search for camera-based semantic scene completion. Expert Syst Appl 243:122885","journal-title":"Expert Syst Appl"},{"key":"1312_CR90","doi-asserted-by":"crossref","unstructured":"Wang X, Lin D, Wan L (2022) Ffnet: frequency fusion network for semantic scene completion. In: AAAI, pp 2550\u20132557","DOI":"10.1609\/aaai.v36i3.20156"},{"key":"1312_CR91","doi-asserted-by":"crossref","unstructured":"Wang F, Sun Q, Zhang D, Tang J (2024) Unleashing network potentials for semantic scene completion. In: CVPR, pp 10314\u201310323","DOI":"10.1109\/CVPR52733.2024.00982"},{"key":"1312_CR92","doi-asserted-by":"crossref","unstructured":"Wang R, Zhang Y, Jia B (2021) Research on the influence of object surface discontinuity on target acoustic scattering characteristics. In: 2021 6th International Conference on Communication, Image and Signal Processing (CCISP), IEEE, pp 345\u2013349.","DOI":"10.1109\/CCISP52774.2021.9639083"},{"key":"1312_CR93","doi-asserted-by":"crossref","unstructured":"Wang F, Zhang D, Zhang H, Tang J, Sun Q (2023) Semantic scene completion with cleaner self. In: CVPR, pp 867\u2013877","DOI":"10.1109\/CVPR52729.2023.00090"},{"key":"1312_CR94","doi-asserted-by":"crossref","unstructured":"Weder S, Sch\u00f6nberger JL, Pollefeys M, Oswald MR (2020) Neuralfusion: online depth fusion in latent space. 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 3161\u20133171","DOI":"10.1109\/CVPR46437.2021.00318"},{"key":"1312_CR95","doi-asserted-by":"crossref","unstructured":"Wolf M, Trentsios P, Kubatzki N, Urbanietz C, Enzner G (2020) Implementing continuous-azimuth binaural sound in unity 3d. In: 2020 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW), IEEE, pp 384\u2013389","DOI":"10.1109\/VRW50115.2020.00083"},{"issue":"9","key":"1312_CR96","doi-asserted-by":"publisher","first-page":"2839","DOI":"10.1016\/j.patcog.2015.03.009","volume":"48","author":"T-T Wong","year":"2015","unstructured":"Wong T-T (2015) Performance evaluation of classification algorithms by k-fold and leave-one-out cross validation. Pattern Recogn 48(9):2839\u20132846","journal-title":"Pattern Recogn"},{"key":"1312_CR97","first-page":"12077","volume":"34","author":"E Xie","year":"2021","unstructured":"Xie E, Wang W, Yu Z, Anandkumar A, Alvarez JM, Luo P (2021) Segformer: simple and efficient design for semantic segmentation with transformers. Adv Neural Inf Process Syst 34:12077\u201312090","journal-title":"Adv Neural Inf Process Syst"},{"key":"1312_CR98","doi-asserted-by":"crossref","unstructured":"Yang S-T, Wang F-E, Peng C-H, Wonka P, Sun M, Chu H-K (2019) Dula-net: A dual-projection network for estimating room layouts from a single rgb panorama. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp 3363\u20133372","DOI":"10.1109\/CVPR.2019.00348"},{"key":"1312_CR99","doi-asserted-by":"crossref","unstructured":"Yao J, Li C, Sun K, Cai Y, Li H, Ouyang W, Li H (2023) Ndc-scene: boost monocular 3d semantic scene completion in normalized device coordinates space. In: 2023 IEEE\/CVF International Conference on Computer Vision (ICCV), IEEE Computer Society, pp 9421\u20139431","DOI":"10.1109\/ICCV51070.2023.00867"},{"key":"1312_CR100","doi-asserted-by":"crossref","unstructured":"Yao Y, Mihalcea R (2022) Modality-specific learning rates for effective multimodal additive late-fusion. In: The Association for Computational Linguistics (ACL), pp 1824\u20131834","DOI":"10.18653\/v1\/2022.findings-acl.143"},{"key":"1312_CR101","doi-asserted-by":"publisher","first-page":"182","DOI":"10.1016\/j.neucom.2018.08.052","volume":"318","author":"L Zhang","year":"2018","unstructured":"Zhang L, Wang L, Zhang X, Shen P, Bennamoun M, Zhu G, Shah SAA, Song J (2018) Semantic scene completion with dense crf from a single depth image. Neurocomputing 318:182\u2013195","journal-title":"Neurocomputing"},{"key":"1312_CR102","doi-asserted-by":"crossref","unstructured":"Zhang P, Liu W, Lei Y, Lu H, Yang X (2019) Cascaded context pyramid for full-resolution 3d semantic scene completion. In: ICCV, pp 7801\u20137810","DOI":"10.1109\/ICCV.2019.00789"},{"key":"1312_CR103","doi-asserted-by":"crossref","unstructured":"Zhang J, Zhao H, Yao A, Chen Y, Zhang L, Liao H (2018) Efficient semantic scene completion network with spatial group convolution. In: ECCV, pp 733\u2013749","DOI":"10.1007\/978-3-030-01258-8_45"},{"key":"1312_CR104","first-page":"2824","volume-title":"European conference on artificial intelligence","author":"M Zhong","year":"2020","unstructured":"Zhong M, Zeng G (2020) Semantic point completion network for 3d semantic scene completion. In: De Giacomo G, Catal\u00e0 A, Montalvo B, Rossi F (eds) European conference on artificial intelligence. IOS Press, Amsterdam, pp 2824\u20132831"}],"container-title":["Virtual Reality"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10055-026-01312-7","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10055-026-01312-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10055-026-01312-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T12:03:24Z","timestamp":1772280204000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10055-026-01312-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,6]]},"references-count":104,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,3]]}},"alternative-id":["1312"],"URL":"https:\/\/doi.org\/10.1007\/s10055-026-01312-7","relation":{},"ISSN":["1434-9957"],"issn-type":[{"value":"1434-9957","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,6]]},"assertion":[{"value":"13 July 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 January 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 February 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"This study involved human participants for subjective evaluation and was approved by the authors\u2019 local institutional ethics committee under reference number ERGO\/FEPS\/99833.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}}],"article-number":"55"}}