{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T08:07:57Z","timestamp":1779350877985,"version":"3.51.4"},"reference-count":106,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T00:00:00Z","timestamp":1773187200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T00:00:00Z","timestamp":1773187200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2026,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    An\n                    <jats:italic>environment acoustic model<\/jats:italic>\n                    represents how sound is transformed by the physical characteristics of an indoor environment, for any given source\/receiver location. Whereas traditional methods for constructing such models assume dense geometry and\/or sound measurements throughout the environment, we explore how to infer room impulse responses (RIRs) based on a sparse set of images and echoes observed in the space, as well as how to choose where to collect these audio-visual observations. Towards that goal, we first introduce a transformer-based method that uses self-attention to build a rich acoustic context, then infers the RIRs of arbitrary query source-receiver locations through cross-attention. Then, motivated by real-world physical constraints in collecting these observations, we further introduce\n                    <jats:italic>active acoustic sampling<\/jats:italic>\n                    , a new task in which a mobile agent jointly constructs the environment acoustic model and spatial occupancy map on-the-fly from sparse audio-visual observations. We train a reinforcement learning (RL) policy that guides agent navigation toward optimal acoustic data sampling positions, rewarding information gain for the full environment model. Evaluating on diverse unseen 3D indoor environments, our method outperforms the state-of-the-art and\u2014in a major departure from traditional methods\u2014generalizes to novel environments in a few-shot manner. Furthermore, when augmented with our active sampling policy, it successfully guides an embodied agent to acoustically informative positions given real-world exploration constraints, outperforming both traditional navigation agents and prior acoustic rendering methods. Project:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"http:\/\/vision.cs.utexas.edu\/projects\/fewShot-RIR\" ext-link-type=\"uri\">http:\/\/vision.cs.utexas.edu\/projects\/fewShot-RIR<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1007\/s11263-026-02767-6","type":"journal-article","created":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T09:11:25Z","timestamp":1773220285000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Sample-efficient Audio-Visual Learning of Scene Acoustics"],"prefix":"10.1007","volume":"134","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-1647-298X","authenticated-orcid":false,"given":"Arjun","family":"Somayazulu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sagnik","family":"Majumder","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Changan","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ziad","family":"Al-Halah","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kristen","family":"Grauman","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,3,11]]},"reference":[{"key":"2767_CR1","doi-asserted-by":"crossref","unstructured":"Afouras, T., Chung, J. S., & Zisserman, A. (2018). The conversation: Deep audio-visual speech enhancement. arXiv preprint arXiv:1804.04121.","DOI":"10.21437\/Interspeech.2018-1400"},{"key":"2767_CR2","doi-asserted-by":"crossref","unstructured":"Afouras, T., Owens, A., Chung, J. S., & Zisserman, A. (2020). Self-supervised learning of audio-visual objects from video. In: European Conference on Computer Vision, pp. 208\u2013224. Springer.","DOI":"10.1007\/978-3-030-58523-5_13"},{"key":"2767_CR3","doi-asserted-by":"publisher","first-page":"943","DOI":"10.1121\/1.382599","volume":"65","author":"J Allen","year":"1979","unstructured":"Allen, J., & Berkley, D. (1979). Image method for efficiently simulating small-room acoustics. The Journal of the Acoustical Society of America, 65, 943\u2013950. https:\/\/doi.org\/10.1121\/1.382599","journal-title":"The Journal of the Acoustical Society of America"},{"key":"2767_CR4","unstructured":"Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., & Munos, R. (2016). Unifying count-based exploration and intrinsic motivation. In: Lee, D., Sugiyama, M., Luxburg, U., Guyon, I., Garnett, R. (eds.) Advances in Neural Information Processing Systems, vol. 29. Curran Associates, Inc."},{"key":"2767_CR5","doi-asserted-by":"publisher","unstructured":"Bilbao, S. (2013). Modeling of complex geometries and boundary conditions in finite difference\/finite volume time domain room acoustics simulation. Trans. Audio, Speech and Lang. Proc.,21(7), 1524\u20131533. https:\/\/doi.org\/10.1109\/TASL.2013.2256897","DOI":"10.1109\/TASL.2013.2256897"},{"issue":"2","key":"2767_CR6","doi-asserted-by":"publisher","first-page":"2312159120","DOI":"10.1073\/pnas.2312159120","volume":"121","author":"N Borrel-Jensen","year":"2024","unstructured":"Borrel-Jensen, N., Goswami, S., Engsig-Karup, A. P., Karniadakis, G. E., & Jeong, C.-H. (2024). Sound propagation in realistic interactive 3d scenes with parameterized sources using deep neural operators. Proceedings of the National Academy of Sciences, 121(2), 2312159120.","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2767_CR7","unstructured":"Bradley, A. J., & Abaid, N. (2024). On fusing active and passive acoustic sensing for simultaneous localization and mapping. https:\/\/arxiv.org\/abs\/2404.13116."},{"key":"2767_CR8","unstructured":"Brunetto, A., Hornauer, S., & Moutarde, F. (2025). NeRAF: 3D Scene Infused Neural Radiance and Acoustic Fields. https:\/\/arxiv.org\/abs\/2405.18213."},{"issue":"6","key":"2767_CR9","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2980179.2982431","volume":"35","author":"C Cao","year":"2016","unstructured":"Cao, C., Ren, Z., Schissler, C., Manocha, D., & Zhou, K. (2016). Interactive sound propagation with bidirectional path tracing. ACM Trans. Graph., 35(6), 1\u201311. https:\/\/doi.org\/10.1145\/2980179.2982431","journal-title":"ACM Trans. Graph."},{"issue":"4","key":"2767_CR10","doi-asserted-by":"publisher","first-page":"44","DOI":"10.1145\/3386569.3392459","volume":"39","author":"CRA Chaitanya","year":"2020","unstructured":"Chaitanya, C. R. A., Raghuvanshi, N., Godin, K. W., Zhang, Z., Nowrouzezahrai, D., & Snyder, J. M. (2020). Directional sources and listeners in interactive sound propagation using reciprocal wave field coding. ACM Trans. Graph., 39(4), 44\u20131. https:\/\/doi.org\/10.1145\/3386569.3392459","journal-title":"ACM Trans. Graph."},{"key":"2767_CR11","doi-asserted-by":"crossref","unstructured":"Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., Song, S., Zeng, A., & Zhang, Y. (2017). Matterport3d: Learning from rgb-d data in indoor environments. arXiv preprint arXiv:1709.06158. Matterport3D license available at http:\/\/kaldir.vc.in.tum.de\/matterport\/MP_TOS.pdf.","DOI":"10.1109\/3DV.2017.00081"},{"key":"2767_CR12","unstructured":"Chen, C., Al-Halah, Z., & Grauman, K. (2020). Semantic audio-visual navigation. 2012.11583"},{"key":"2767_CR13","doi-asserted-by":"crossref","unstructured":"Chen, C., Al-Halah, Z., & Grauman, K. (2021). Semantic Audio-Visual Navigation.","DOI":"10.1109\/CVPR46437.2021.01526"},{"key":"2767_CR14","doi-asserted-by":"crossref","unstructured":"Chen, C., Gao, R., Calamia, P., & Grauman, K. (2022). Visual acoustic matching. arXiv preprint arXiv:2202.06875.","DOI":"10.1109\/CVPR52688.2022.01829"},{"key":"2767_CR15","doi-asserted-by":"crossref","unstructured":"Chen, Z., Gebru, I. D., Richardt, C., Kumar, A., Laney, W., Owens, A., & Richard, A. (2024). Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark. https:\/\/arxiv.org\/abs\/2403.18821.","DOI":"10.1109\/CVPR52733.2024.02067"},{"key":"2767_CR16","unstructured":"Chen, T., Gupta, S., & Gupta, A. (2019). Learning Exploration Policies for Navigation."},{"key":"2767_CR17","doi-asserted-by":"crossref","unstructured":"Chen, C., Jain, U., Schissler, C., Gari, S. V. A., Al-Halah, Z., Ithapu, V. K., Robinson, P., & Grauman, K. (2020) Soundspaces: Audio-visual navigation in 3d environments. In: European Conference on Computer Vision, pp. 17\u201336. Springer.","DOI":"10.1007\/978-3-030-58539-6_2"},{"key":"2767_CR18","unstructured":"Chen, C., Majumder, S., Al-Halah, Z., Gao, R., Ramakrishnan, S. K., & Grauman, K. (2020). Audio-visual waypoints for navigation. CoRR abs\/2008.09622."},{"key":"2767_CR19","unstructured":"Chen, C., Schissler, C., Garg, S., Kobernik, P., Clegg, A., Calamia, P., Batra, D., Robinson, P. W., & Grauman, K. (2023). SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning."},{"key":"2767_CR20","doi-asserted-by":"crossref","unstructured":"Chen, C., Sun, W., Harwath, D., & Grauman, K. (2023). Learning audio-visual dereverberation. arXiv preprint arXiv:2106.07732 (2021).","DOI":"10.1109\/ICASSP49357.2023.10095818"},{"key":"2767_CR21","doi-asserted-by":"crossref","unstructured":"Christensen, J. H., Hornauer, S., & Stella, X. Y. (2020). Batvision: Learning to see 3d spatial layout with two ears. In: 2020 IEEE International Conference on Robotics and Automation (ICRA), pp. 1581\u20131587 IEEE.","DOI":"10.1109\/ICRA40945.2020.9196934"},{"key":"2767_CR22","unstructured":"Dean, V., Tulsiani, S., & Gupta, A. (2020). See, hear, explore: Curiosity via audio-visual association. In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M.F., Lin, H. (eds.) Advances in Neural Information Processing Systems, vol. 33, pp. 14961\u201314972. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/ab6b331e94c28169d15cca0cb3bbc73e-Paper.pdf."},{"issue":"10","key":"2767_CR23","doi-asserted-by":"publisher","first-page":"1681","DOI":"10.1109\/TASLP.2016.2577502","volume":"24","author":"J Eaton","year":"2016","unstructured":"Eaton, J., Gaubitch, N. D., Moore, A. H., & Naylor, P. A. (2016). Estimation of room acoustic parameters: The ace challenge. IEEE\/ACM Transactions on Audio, Speech, and Language Processing, 24(10), 1681\u20131693. https:\/\/doi.org\/10.1109\/TASLP.2016.2577502","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"2767_CR24","doi-asserted-by":"crossref","unstructured":"Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., Freeman, W. T., & Rubinstein, M. (2018). Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. arXiv preprint arXiv:1804.03619.","DOI":"10.1145\/3197517.3201357"},{"key":"2767_CR25","unstructured":"Farina, A.: (2000). Simultaneous measurement of impulse response and distortion with a swept-sine technique. Journal of the Audio Engineering Society."},{"key":"2767_CR26","doi-asserted-by":"publisher","unstructured":"Feng, Y., Khodayi-mehr, R., Kantaros, Y., Calkins, L., & Zavlanos, M. M. (2018). Active acoustic impedance mapping using mobile robots. In: 2018 IEEE Conference on Decision and Control (CDC), pp. 3910\u20133915. https:\/\/doi.org\/10.1109\/CDC.2018.8618924.","DOI":"10.1109\/CDC.2018.8618924"},{"issue":"2","key":"2767_CR27","first-page":"739","volume":"115","author":"T Funkhouser","year":"2003","unstructured":"Funkhouser, T., Tsingos, N., Carlbom, I., Elko, G., Sondhi, M., West, J. E., Pingali, G., Min, P., & Ngan, A. (2003). A Beam Tracing Method for Interactive Architectural Acoustics, 115(2), 739\u2013756.","journal-title":"A Beam Tracing Method for Interactive Architectural Acoustics"},{"key":"2767_CR28","doi-asserted-by":"publisher","unstructured":"Gamper, H., & Tashev, I. J. (2018). Blind reverberation time estimation using a convolutional neural network. In: 2018 16th International Workshop on Acoustic Signal Enhancement (IWAENC), pp. 136\u2013140 https:\/\/doi.org\/10.1109\/IWAENC.2018.8521241.","DOI":"10.1109\/IWAENC.2018.8521241"},{"key":"2767_CR29","doi-asserted-by":"crossref","unstructured":"Gan, C., Zhang, Y., Wu, J., Gong, B., & Tenenbaum, J. B. (2020). Look, listen, and act: Towards audio-visual embodied navigation. In: 2020 IEEE International Conference on Robotics and Automation (ICRA), pp. 9701\u20139707 IEEE.","DOI":"10.1109\/ICRA40945.2020.9197008"},{"key":"2767_CR30","doi-asserted-by":"crossref","unstructured":"Gao, R., & Grauman, K. (2019). 2.5 d visual sound. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 324\u2013333.","DOI":"10.1109\/CVPR.2019.00041"},{"key":"2767_CR31","doi-asserted-by":"crossref","unstructured":"Gao, R., Chen, C., Al-Halab, Z., Schissler, C., & Grauman, K. (2020). Visualechoes: Spatial image representation learning through echolocation. In: ECCV.","DOI":"10.1007\/978-3-030-58545-7_38"},{"key":"2767_CR32","unstructured":"Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. Generative adversarial nets. In: Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N., Weinberger, K.Q. (eds.) Advances in Neural Information Processing Systems, vol. 27. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2014\/file\/f033ed80deb0234979a61f95710dbe25-Paper.pdf"},{"key":"2767_CR33","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S. & Sun, J. (2016). Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770\u2013778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"2767_CR34","unstructured":"Holters, M., Corbach, T., & Z\u00f6lzer, U. (2009). IMPULSE RESPONSE MEASUREMENT TECHNIQUES AND THEIR APPLICABILITY IN THE REAL WORLD."},{"issue":"2","key":"2767_CR35","doi-asserted-by":"publisher","first-page":"117","DOI":"10.1109\/TETCI.2017.2784878","volume":"2","author":"J-C Hou","year":"2018","unstructured":"Hou, J.-C., Wang, S.-S., Lai, Y.-H., Tsao, Y., Chang, H.-W., & Wang, H.-M. (2018). Audio-visual speech enhancement using multimodal deep convolutional neural networks. IEEE Transactions on Emerging Topics in Computational Intelligence, 2(2), 117\u2013128.","journal-title":"IEEE Transactions on Emerging Topics in Computational Intelligence"},{"key":"2767_CR36","doi-asserted-by":"crossref","unstructured":"Hu, X., Purushwalkam, S., Harwath, D., & Grauman, K. (2023). Learning to Map Efficiently by Active Echolocation.","DOI":"10.1109\/IROS55552.2023.10341664"},{"key":"2767_CR37","first-page":"10077","volume":"33","author":"D Hu","year":"2020","unstructured":"Hu, D., Qian, R., Jiang, M., Tan, X., Wen, S., Ding, E., Lin, W., & Dou, D. (2020). Discriminative sounding objects localization via self-supervised audiovisual matching. Advances in Neural Information Processing Systems, 33, 10077\u201310087.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2767_CR38","unstructured":"Ioffe, S., & Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: Bach, F., Blei, D. (eds.) Proceedings of the 32nd International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 37, pp. 448\u2013456. PMLR, Lille, France. https:\/\/proceedings.mlr.press\/v37\/ioffe15.html."},{"key":"2767_CR39","doi-asserted-by":"crossref","unstructured":"Jiang, H., Murdock, C., & Ithapu, V. K. (2022). Egocentric deep multi-channel audio-visual active speaker localization. arXiv preprint arXiv:2201.01928.","DOI":"10.1109\/CVPR52688.2022.01029"},{"key":"2767_CR40","first-page":"3255","volume":"2024","author":"L Kelley","year":"2024","unstructured":"Kelley, L., Di Carlo, D., Nugraha, A. A., Fontaine, M., Bando, Y., & Yoshii, K. (2024). Rir-in-a-box: Estimating room acoustics from 3d mesh data through shoebox approximation. Interspeech, 2024, 3255\u20133259.","journal-title":"Interspeech"},{"key":"2767_CR41","unstructured":"Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980."},{"key":"2767_CR42","doi-asserted-by":"publisher","unstructured":"Klein, F., Neidhardt, A., & Seipel, M. (2019). Real-time estimation of reverberation time for selection of suitable binaural room impulse responses. Audio for Virtual, Augmented and Mixed Realities: Proceedings of ICSA 2019; 5th International Conference on Spatial Audio; September 26th to 28th, 2019, Ilmenau, Germany, 145\u2013150 https:\/\/doi.org\/10.22032\/dbt.39968.","DOI":"10.22032\/dbt.39968"},{"key":"2767_CR43","doi-asserted-by":"publisher","unstructured":"Kon, H., & Koike, H. (2019). Estimation of late reverberation characteristics from a single two-dimensional environmental image using convolutional neural networks. Journal of the Audio Engineering Society,67(7\/8), 540\u2013548. https:\/\/doi.org\/10.17743\/jaes.2018.0069","DOI":"10.17743\/jaes.2018.0069"},{"key":"2767_CR44","first-page":"359","volume":"87","author":"K Lebart","year":"2001","unstructured":"Lebart, K., Boucher, J.-M., & Denbigh, P. (2001). A new method based on spectral subtraction for speech dereverberation. Acta Acustica united with Acustica, 87, 359\u2013366.","journal-title":"Acta Acustica united with Acustica"},{"key":"2767_CR45","unstructured":"Liang, S., Huang, C., Tian, Y., Kumar, A., & Xu, C. (2023). Neural Acoustic Context Field: Rendering Realistic Room Impulse Response With Neural Fields."},{"key":"2767_CR46","first-page":"37472","volume":"36","author":"S Liang","year":"2023","unstructured":"Liang, S., Huang, C., Tian, Y., Kumar, A., & Xu, C. (2023). AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene. Synthesis, 36, 37472\u201337490. arxiv.org\/abs\/2302.02088.","journal-title":"Synthesis"},{"key":"2767_CR47","unstructured":"Liu, X., Paul, S., Chatterjee, M., & Cherian, A. (2023). CAVEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy Environments. https:\/\/arxiv.org\/abs\/2306.04047."},{"key":"2767_CR48","first-page":"3165","volume":"35","author":"A Luo","year":"2023","unstructured":"Luo, A., Du, Y., Tarr, M. J., Tenenbaum, J. B., Torralba, A., & Gan, C. (2023). Learning Neural Acoustic Fields, 35, 3165\u20133177.","journal-title":"Learning Neural Acoustic Fields"},{"key":"2767_CR49","doi-asserted-by":"publisher","unstructured":"Mack, W., Deng, S., & Habets, E. A. P. (2020) Single-channel blind direct-to-reverberation ratio estimation using masking. In: Meng, H., Xu, B., Zheng, T.F. (eds.) Interspeech 2020, 21st Annual Conference of the International Speech Communication Association, Virtual Event, Shanghai, China, 25-29 October, pp. 5066\u20135070. ISCA. https:\/\/doi.org\/10.21437\/Interspeech.2020-2171 .","DOI":"10.21437\/Interspeech.2020-2171"},{"key":"2767_CR50","doi-asserted-by":"crossref","unstructured":"Majumder, S., Al-Halah, Z., & Grauman, K. (2021). Move2hear: Active audio-visual source separation. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 275\u2013285.","DOI":"10.1109\/ICCV48922.2021.00034"},{"key":"2767_CR51","doi-asserted-by":"crossref","unstructured":"Majumder, S., Al-Halah, Z., & Grauman, K. (2022). Active audio-visual separation of dynamic sound sources. arXiv preprint arXiv:2202.00850.","DOI":"10.1007\/978-3-031-19842-7_32"},{"key":"2767_CR52","doi-asserted-by":"crossref","unstructured":"Majumder, S., Jiang, H., Moulon, P., Henderson, E., Calamia, P., Grauman, K., & Ithapu, V. K. (2023). Chat2Map: Efficient Scene Mapping from Multi-Ego Conversations.","DOI":"10.1109\/CVPR52729.2023.01017"},{"key":"2767_CR53","first-page":"2522","volume":"35","author":"S Majumder","year":"2022","unstructured":"Majumder, S., Chen, C., Al-Halah, Z., & Grauman, K. (2022). Few-Shot Audio-Visual Learning of Environment Acoustics, 35, 2522\u20132536.","journal-title":"Few-Shot Audio-Visual Learning of Environment Acoustics"},{"issue":"4","key":"2767_CR54","doi-asserted-by":"publisher","first-page":"495","DOI":"10.1109\/TVCG.2014.38","volume":"20","author":"R Mehra","year":"2014","unstructured":"Mehra, R., Antani, L., Kim, S., & Manocha, D. (2014). Source and listener directivity for interactive wave-based sound propagation. IEEE Transactions on Visualization and Computer Graphics, 20(4), 495\u2013503. https:\/\/doi.org\/10.1109\/TVCG.2014.38","journal-title":"IEEE Transactions on Visualization and Computer Graphics"},{"issue":"4","key":"2767_CR55","doi-asserted-by":"publisher","first-page":"495","DOI":"10.1109\/TVCG.2014.38","volume":"20","author":"R Mehra","year":"2014","unstructured":"Mehra, R., Antani, L., Kim, S., & Manocha, D. (2014). Source and listener directivity for interactive wave-based sound propagation. IEEE Transactions on Visualization and Computer Graphics, 20(4), 495\u2013503. https:\/\/doi.org\/10.1109\/TVCG.2014.38","journal-title":"IEEE Transactions on Visualization and Computer Graphics"},{"key":"2767_CR56","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3066303","volume-title":"An overview of deep-learning-based audio-visual speech enhancement and separation","author":"D Michelsanti","year":"2021","unstructured":"Michelsanti, D., Tan, Z.-H., Zhang, S.-X., Xu, Y., Yu, M., Yu, D., & Jensen, J. (2021). An overview of deep-learning-based audio-visual speech enhancement and separation. IEEE\/ACM Transactions on Audio: Speech, and Language Processing."},{"key":"2767_CR57","doi-asserted-by":"crossref","unstructured":"Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., & Ng, R. (2020). Nerf: Representing scenes as neural radiance fields for view synthesis. In: European Conference on Computer Vision, pp. 405\u2013421. Springer.","DOI":"10.1007\/978-3-030-58452-8_24"},{"issue":"2","key":"2767_CR58","doi-asserted-by":"publisher","first-page":"55","DOI":"10.1109\/MSP.2007.323264","volume":"24","author":"D Murphy","year":"2007","unstructured":"Murphy, D., Kelloniemi, A., Mullen, J., & Shelley, S. (2007). Acoustic modeling using the digital waveguide mesh. IEEE Signal Processing Magazine, 24(2), 55\u201366. https:\/\/doi.org\/10.1109\/MSP.2007.323264","journal-title":"IEEE Signal Processing Magazine"},{"key":"2767_CR59","unstructured":"Nair, V., & Hinton, G. E. (2010). Rectified linear units improve restricted boltzmann machines. In: ICML, pp. 807\u2013814. https:\/\/icml.cc\/Conferences\/2010\/papers\/432.pdf"},{"key":"2767_CR60","doi-asserted-by":"crossref","unstructured":"Owens, A., & Efros, A. A. (2018). Audio-visual scene analysis with self-supervised multisensory features. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 631\u2013648.","DOI":"10.1007\/978-3-030-01231-1_39"},{"key":"2767_CR61","doi-asserted-by":"publisher","unstructured":"Piczak, K. J. (2015). ESC: Dataset for Environmental Sound Classification. In: Proceedings of the 23rd Annual ACM Conference on Multimedia, pp. 1015\u20131018. ACM Press. https:\/\/doi.org\/10.1145\/2733373.2806390.","DOI":"10.1145\/2733373.2806390"},{"key":"2767_CR62","doi-asserted-by":"crossref","unstructured":"Purushwalkam, S., Gari, S. V. A., Ithapu, V. K., Schissler, C., Robinson, P., Gupta, A., & Grauman, K. (2020). Audio-visual floorplan reconstruction., arXiv preprint arXiv:2012.15470.","DOI":"10.1109\/ICCV48922.2021.00122"},{"key":"2767_CR63","doi-asserted-by":"crossref","unstructured":"Purushwalkam, S., Gari, S. V. A., Ithapu, V. K., Schissler, C., Robinson, P., Gupta, A., & Grauman, K. (2021). Audio-visual floorplan reconstruction. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 1183\u20131192.","DOI":"10.1109\/ICCV48922.2021.00122"},{"key":"2767_CR64","doi-asserted-by":"crossref","unstructured":"Rachavarapu, K. K., Aakanksha, Sundaresha, V., & Rajagopalan, A. N. (2021). Localize to binauralize: Audio spatialization from visual sound source localization. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), pp. 1930\u20131939.","DOI":"10.1109\/ICCV48922.2021.00194"},{"issue":"4","key":"2767_CR65","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2601097.2601184","volume":"33","author":"N Raghuvanshi","year":"2014","unstructured":"Raghuvanshi, N., & Snyder, J. (2014). Parametric wave field coding for precomputed sound propagation. ACM Trans. Graph., 33(4), 1\u201311. https:\/\/doi.org\/10.1145\/2601097.2601184","journal-title":"ACM Trans. Graph."},{"issue":"4","key":"2767_CR66","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2601097.2601184","volume":"33","author":"N Raghuvanshi","year":"2014","unstructured":"Raghuvanshi, N., & Snyder, J. (2014). Parametric Wave Field Coding for Precomputed Sound Propagation. ACM Transactions on Graphics (TOG), 33(4), 1\u201311.","journal-title":"ACM Transactions on Graphics (TOG)"},{"issue":"4","key":"2767_CR67","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3197517.3201339","volume":"37","author":"N Raghuvanshi","year":"2018","unstructured":"Raghuvanshi, N., & Snyder, J. (2018). Parametric directional coding for precomputed sound propagation. ACM Trans. Graph., 37(4), 1\u201314. https:\/\/doi.org\/10.1145\/3197517.3201339","journal-title":"ACM Trans. Graph."},{"issue":"1","key":"2767_CR68","volume":"2872","author":"SK Ramakrishnan","year":"2020","unstructured":"Ramakrishnan, S. K., Jayaraman, D., & Grauman, K. (2020). An Exploration of Embodied Visual Exploration, 2872(1), Article 060029.","journal-title":"An Exploration of Embodied Visual Exploration"},{"key":"2767_CR69","doi-asserted-by":"crossref","unstructured":"Ratnarajah, A., & Manocha, D. (2024). Listen2Scene: Interactive material-aware binaural sound propagation for reconstructed 3D scenes.","DOI":"10.1109\/VR58804.2024.00048"},{"key":"2767_CR70","doi-asserted-by":"crossref","unstructured":"Ratnarajah, A., Tang, Z., & Manocha, D. (2020). Ir-gan: Room impulse response generator for far-field speech recognition. arXiv preprint arXiv:2010.13219.","DOI":"10.21437\/Interspeech.2021-230"},{"key":"2767_CR71","doi-asserted-by":"crossref","unstructured":"Ratnarajah, A., Tang, Z., & Manocha, D. (2021). Ts-rir: Translated synthetic room impulse responses for speech augmentation. arXiv preprint arXiv:2103.16804.","DOI":"10.1109\/ASRU51503.2021.9688304"},{"key":"2767_CR72","doi-asserted-by":"crossref","unstructured":"Ratnarajah, A., Tang, Z., Aralikatti, R. C., & Manocha, D. (2022). Mesh2ir: Neural acoustic impulse response generator for complex 3d scenes. arXiv preprint arXiv:2205.09248","DOI":"10.1145\/3503161.3548253"},{"key":"2767_CR73","doi-asserted-by":"crossref","unstructured":"Ratnarajah, A., Zhang, S.-X., Yu, M., Tang, Z., Manocha, D., & Yu, D. (2022) Fast-rir: Fast neural diffuse room impulse response generator. In: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 571\u2013575. IEEE.","DOI":"10.1109\/ICASSP43922.2022.9747846"},{"key":"2767_CR74","doi-asserted-by":"publisher","unstructured":"Ratnarajah, A., Zhang, S. X., Yu, M., Tang, Z., Manocha, D., & Yu, D. (2022). Fast-rir: Fast neural diffuse room impulse response generator. In: ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 571\u2013575. https:\/\/doi.org\/10.1109\/ICASSP43922.2022.9747846.","DOI":"10.1109\/ICASSP43922.2022.9747846"},{"key":"2767_CR75","unstructured":"Remaggi, L., Kim, H., Jackson, P. J. B., & Hilton, A. (2019). Reproducing real world acoustics in virtual reality using spherical cameras. Journal of the Audio Engineering Society."},{"key":"2767_CR76","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P. & Brox, T. (2015). U-net: Convolutional networks for biomedical image segmentation. In: International Conference on Medical Image Computing and Computer-assisted Intervention, pp. 234\u2013241. Springer.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"2767_CR77","doi-asserted-by":"publisher","first-page":"1788","DOI":"10.1109\/TASLP.2020.3000593","volume":"28","author":"M Sadeghi","year":"2020","unstructured":"Sadeghi, M., Leglaive, S., Alameda-Pineda, X., Girin, L., & Horaud, R. (2020). Audio-visual speech enhancement using conditional variational auto-encoders. IEEE\/ACM Transactions on Audio, Speech, and Language Processing, 28, 1788\u20131800.","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"2767_CR78","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2021.104009","volume":"150","author":"U Saqib","year":"2022","unstructured":"Saqib, U., & Jensen, J. R. (2022). A framework for spatial map generation using acoustic echoes for robotic platforms. Robotics and Autonomous Systems, 150, Article 104009. https:\/\/doi.org\/10.1016\/j.robot.2021.104009","journal-title":"Robotics and Autonomous Systems"},{"key":"2767_CR79","unstructured":"Savinov, N., Raichuk, A., Marinier, R., Vincent, D., Pollefeys, M., Lillicrap, T., & Gelly, S. (2019). Episodic curiosity through reachability., In: International Conference on Learning Representations (ICLR)."},{"issue":"2","key":"2767_CR80","doi-asserted-by":"publisher","first-page":"708","DOI":"10.1121\/1.4926438","volume":"138","author":"L Savioja","year":"2015","unstructured":"Savioja, L., & Svensson, U. P. (2015). Overview of geometrical room acoustic modeling techniques. The Journal of the Acoustical Society of America, 138(2), 708\u201330.","journal-title":"The Journal of the Acoustical Society of America"},{"key":"2767_CR81","doi-asserted-by":"crossref","unstructured":"Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Parikh, D., & Batra, D. (2019). Habitat: A Platform for Embodied AI Research. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV).","DOI":"10.1109\/ICCV.2019.00943"},{"key":"2767_CR82","unstructured":"Schulman, J., Moritz, P., Levine, S., Jordan, M., & Abbeel, P. (2015). High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438."},{"key":"2767_CR83","doi-asserted-by":"crossref","unstructured":"Senocak, A., Oh, T.-H., Kim, J., Yang, M.-H., & Kweon, I. S. (2018). Learning to Localize Sound Source in Visual Scenes. https:\/\/arxiv.org\/abs\/1803.03849.","DOI":"10.1109\/CVPR.2018.00458"},{"key":"2767_CR84","doi-asserted-by":"crossref","unstructured":"Senocak, A., Ryu, H., Kim, J., Oh, T.-H., Pfister, H., & Chung, J. S. (2023). Sound Source Localization is All about Cross-Modal Alignment. https:\/\/arxiv.org\/abs\/2309.10724.","DOI":"10.1109\/ICCV51070.2023.00715"},{"key":"2767_CR85","unstructured":"Senocak, A., Ryu, H., Kim, J., Oh, T.-H., Pfister, H., & Chung, J. S. (2024). Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment. https:\/\/arxiv.org\/abs\/2407.13676"},{"key":"2767_CR86","unstructured":"Si, C., Wu, Q., Amballa, C., & Choudhury, R. R. (2025). Explicit Context-Driven Neural Acoustic Modeling for High-Fidelity RIR Generation. https:\/\/arxiv.org\/abs\/2509.15210."},{"key":"2767_CR87","doi-asserted-by":"crossref","unstructured":"Simpson, A. J. R., Roma, G., & Plumbley, M. D. (2015). Deep Karaoke: Extracting Vocals from Musical Mixtures Using a Convolutional Deep Neural Network. https:\/\/arxiv.org\/abs\/1504.04658.","DOI":"10.1007\/978-3-319-22482-4_50"},{"key":"2767_CR88","doi-asserted-by":"crossref","unstructured":"Singh, N., Mentch, J., Ng, J., Beveridge, & M., Drori, I. (2021). Image2reverb: Cross-modal reverb impulse response synthesis. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 286\u201329.","DOI":"10.1109\/ICCV48922.2021.00035"},{"key":"2767_CR89","first-page":"7462","volume":"33","author":"V Sitzmann","year":"2020","unstructured":"Sitzmann, V., Martel, J., Bergman, A., Lindell, D., & Wetzstein, G. (2020). Implicit neural representations with periodic activation functions. Advances in Neural Information Processing Systems, 33, 7462\u20137473.","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"56","key":"2767_CR90","first-page":"1929","volume":"15","author":"N Srivastava","year":"2014","unstructured":"Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(56), 1929\u20131958.","journal-title":"Journal of Machine Learning Research"},{"key":"2767_CR91","unstructured":"Stan, G.-B., Embrechts, J.-J., & Archambeau, D. (2002). Comparison of different impulse response measurement techniques. Journal of the Audio Engineering Society,50(4), 249\u2013262."},{"issue":"8","key":"2767_CR92","doi-asserted-by":"publisher","first-page":"1309","DOI":"10.1016\/j.jcss.2007.08.009","volume":"74","author":"AL Strehl","year":"2008","unstructured":"Strehl, A. L., & Littman, M. L. (2008). An analysis of model-based Interval Estimation for Markov Decision Processes. Journal of Computer and System Sciences, 74(8), 1309\u20131331. https:\/\/doi.org\/10.1016\/j.jcss.2007.08.009","journal-title":"Journal of Computer and System Sciences"},{"key":"2767_CR93","doi-asserted-by":"crossref","unstructured":"Su, K., Chen, M., & Shlizerman, E. Inras: Implicit neural representation for audio scenes. In: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A. (eds.) Advances in Neural Information Processing Systems, vol. 35, pp. 8144\u20138158. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/35d5ad984cc0ddd84c6f1c177a2066e5-Paper-Conference.pdf.","DOI":"10.52202\/068431-0591"},{"key":"2767_CR94","doi-asserted-by":"crossref","unstructured":"Sun, Y., Wang, X., & Tang, X. (2015) Deeply learned face representations are sparse, selective, and robust. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2892\u20132900.","DOI":"10.1109\/CVPR.2015.7298907"},{"key":"2767_CR95","unstructured":"V\u00e4lim\u00e4ki, V., Parker, J., Savioja, L., Smith, J., & Abel, J. (2016). More than 50 years of artificial reverberation. In: Goetze, S., Spriet, A. (eds.) Proc. 60th International Conference of the Audio Engineering Society. Audio Engineering Society, United States. AES International Conference on Dereverberation and Reverberation of Audio, Music, and Speech, DREAMS ; Conference date: 03-02-2016 Through 05-02-2016."},{"key":"2767_CR96","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, \u0141, & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems,30."},{"issue":"3","key":"2767_CR97","doi-asserted-by":"publisher","first-page":"829","DOI":"10.1109\/JOE.2020.3033036","volume":"46","author":"Y Wang","year":"2021","unstructured":"Wang, Y., Ji, Y., Woo, H., Tamura, Y., Tsuchiya, H., Yamashita, A., & Asama, H. (2021). Acoustic camera-based pose graph slam for dense 3-d mapping in underwater environments. IEEE Journal of Oceanic Engineering, 46(3), 829\u2013847. https:\/\/doi.org\/10.1109\/JOE.2020.3033036","journal-title":"IEEE Journal of Oceanic Engineering"},{"issue":"3","key":"2767_CR98","doi-asserted-by":"publisher","first-page":"1150","DOI":"10.3390\/app11031150","volume":"11","author":"S Werner","year":"2021","unstructured":"Werner, S., Klein, F., Neidhardt, A., Sloma, U., Schneiderwind, C., & Brandenburg, K. (2021). Creation of auditory augmented reality using a position-dynamic binaural synthesis system-technical components, psychoacoustic needs, and perceptual evaluation. Applied Sciences, 11(3), 1150. https:\/\/doi.org\/10.3390\/app11031150","journal-title":"Applied Sciences"},{"key":"2767_CR99","unstructured":"Wijmans, E., Kadian, A., Morcos, A., Lee, S., Essa, I., Parikh, D., Savva, M., & Batra, D. (2019). Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames. arXiv preprint arXiv:1911.00357."},{"issue":"2","key":"2767_CR100","doi-asserted-by":"publisher","first-page":"255","DOI":"10.1109\/TASLP.2018.2877894","volume":"27","author":"F Xiong","year":"2019","unstructured":"Xiong, F., Goetze, S., Kollmeier, B., & Meyer, B. T. (2019). Joint estimation of reverberation time and early-to-late reverberation ratio from single-channel speech signals. IEEE\/ACM Transactions on Audio, Speech, and Language Processing, 27(2), 255\u2013267. https:\/\/doi.org\/10.1109\/TASLP.2018.2877894","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"2767_CR101","doi-asserted-by":"crossref","unstructured":"Xu, X., Zhou, H., Liu, Z., Dai, B., Wang, X., & Lin, D. (2021) Visually informed binaural audio generation without binaural audios. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 15485\u201315494.","DOI":"10.1109\/CVPR46437.2021.01523"},{"key":"2767_CR102","doi-asserted-by":"publisher","unstructured":"Yu, Y., Chen, C., Cao, L., Yang, F., & Sun, F. Measuring Acoustics with Collaborative Multiple Agents. In: Elkind, E. (ed.) Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, IJCAI-23, pp. 335\u2013343. International Joint Conferences on Artificial Intelligence Organization. https:\/\/doi.org\/10.24963\/ijcai.2023\/38.","DOI":"10.24963\/ijcai.2023\/38"},{"key":"2767_CR103","unstructured":"Yu, Y., Huang, W., Sun, F., Chen, C., Wang, Y., & Liu, X. (2022). Sound Adversarial Audio-Visual Navigation."},{"key":"2767_CR104","doi-asserted-by":"crossref","unstructured":"Zhao, H., Gan, C., Ma, W.-C., & Torralba, A. (2019). The sound of motions. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 1735\u20131744.","DOI":"10.1109\/ICCV.2019.00182"},{"key":"2767_CR105","doi-asserted-by":"crossref","unstructured":"Zhou, H., Liu, Z., Xu, X., Luo, P., & Wang, X. (2019). Vision-infused deep audio inpainting. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 283\u2013292.","DOI":"10.1109\/ICCV.2019.00037"},{"key":"2767_CR106","doi-asserted-by":"crossref","unstructured":"Zhou, H., Xu, X., Lin, D., Wang, X., & Liu, Z. (2020) Sep-stereo: Visually guided stereophonic audio generation by associating source separation. CoRR abs\/2007.09902.","DOI":"10.1007\/978-3-030-58610-2_4"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02767-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-026-02767-6","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02767-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T07:33:53Z","timestamp":1779348833000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-026-02767-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,11]]},"references-count":106,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4]]}},"alternative-id":["2767"],"URL":"https:\/\/doi.org\/10.1007\/s11263-026-02767-6","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,11]]},"assertion":[{"value":"16 April 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 January 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 March 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"193"}}