{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T11:46:19Z","timestamp":1780487179522,"version":"3.54.1"},"reference-count":29,"publisher":"Springer Science and Business Media LLC","issue":"10-12","license":[{"start":{"date-parts":[[2020,7,31]],"date-time":"2020-07-31T00:00:00Z","timestamp":1596153600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,7,31]],"date-time":"2020-07-31T00:00:00Z","timestamp":1596153600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"JST ACCEL","award":["JPMJAC1602"],"award-info":[{"award-number":["JPMJAC1602"]}]},{"name":"JST ACCEL","award":["JPMJAC1602"],"award-info":[{"award-number":["JPMJAC1602"]}]},{"name":"JST-Mirai Program","award":["JPMJMI19B2"],"award-info":[{"award-number":["JPMJMI19B2"]}]},{"name":"JSPS KAKENHI","award":["JP19H01129"],"award-info":[{"award-number":["JP19H01129"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis Comput"],"published-print":{"date-parts":[[2020,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We present a novel concept<jats:italic>audio\u2013visual object removal<\/jats:italic>in 360-degree videos, in which a target object in a 360-degree video is removed in both the visual and auditory domains synchronously. Previous methods have solely focused on the visual aspect of object removal using video inpainting techniques, resulting in videos with unreasonable remaining sounds corresponding to the removed objects. We propose a solution which incorporates direction acquired during the video inpainting process into the audio removal process. More specifically, our method identifies the sound corresponding to the visually tracked target object and then synthesizes a three-dimensional sound field by subtracting the identified sound from the input 360-degree video. We conducted a user study showing that our multi-modal object removal supporting both visual and auditory domains could significantly improve the virtual reality experience, and our method could generate sufficiently synchronous, natural and satisfactory 360-degree videos.<\/jats:p>","DOI":"10.1007\/s00371-020-01918-1","type":"journal-article","created":{"date-parts":[[2020,7,31]],"date-time":"2020-07-31T01:02:53Z","timestamp":1596157373000},"page":"2117-2128","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Audio\u2013visual object removal in 360-degree videos"],"prefix":"10.1007","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8219-7470","authenticated-orcid":false,"given":"Ryo","family":"Shimamura","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6892-3122","authenticated-orcid":false,"given":"Qi","family":"Feng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3978-1444","authenticated-orcid":false,"given":"Yuki","family":"Koyama","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3181-4894","authenticated-orcid":false,"given":"Takayuki","family":"Nakatsuka","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6506-2796","authenticated-orcid":false,"given":"Satoru","family":"Fukayama","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3085-7446","authenticated-orcid":false,"given":"Masahiro","family":"Hamasaki","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1167-0977","authenticated-orcid":false,"given":"Masataka","family":"Goto","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8859-6539","authenticated-orcid":false,"given":"Shigeo","family":"Morishima","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,7,31]]},"reference":[{"key":"1918_CR1","doi-asserted-by":"crossref","unstructured":"Akyazi, P., Frossard, P.: Graph-based inpainting of disocclusion holes for zooming in 3d scenes. In: 2018 26th European Signal Processing Conference (EUSIPCO), pp. 867\u2013871. IEEE (2018)","DOI":"10.23919\/EUSIPCO.2018.8553205"},{"key":"1918_CR2","doi-asserted-by":"crossref","unstructured":"Bertalmio, M., Bertozzi, A.L., Sapiro, G.: Navier\u2013Stokes, fluid dynamics, and image and video inpainting. In: Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. CVPR 2001, vol.\u00a01, pp. I\u2013I. IEEE (2001)","DOI":"10.1109\/CVPR.2001.990497"},{"key":"1918_CR3","doi-asserted-by":"crossref","unstructured":"Bertalm\u00edo, M., Caselles, V., Haro, G., Sapiro, G.: Pde-based image and surface inpainting. In: Handbook of Mathematical Models in Computer Vision, pp. 33\u201361. Springer (2006)","DOI":"10.1007\/0-387-28831-7_3"},{"key":"1918_CR4","unstructured":"Facebook: How do I upload a 360 video on Facebook?\u2014Facebook Help Center. Retrieved November 19, 2019 from https:\/\/www.facebook.com\/help\/828417127257368"},{"key":"1918_CR5","doi-asserted-by":"crossref","unstructured":"Feng, W., Guan, N., Li, Y., Zhang, X., Luo, Z.: Audio visual speech recognition with multimodal recurrent neural networks. In: 2017 International Joint Conference on Neural Networks (IJCNN), pp. 681\u2013688. IEEE. (2017)","DOI":"10.1109\/IJCNN.2017.7965918"},{"issue":"1","key":"1918_CR6","first-page":"2","volume":"21","author":"MA Gerzon","year":"1973","unstructured":"Gerzon, M.A.: Periphony: with-height sound reproduction. J. Audio Eng. Soc. 21(1), 2\u201310 (1973)","journal-title":"J. Audio Eng. Soc."},{"key":"1918_CR7","unstructured":"Google: Upload 360-degree videos\u2014YouTube Help. Retrieved November 19, 2019 from https:\/\/support.google.com\/youtube\/answer\/6178631"},{"key":"1918_CR8","unstructured":"Hershey, J.R., Casey, M.: Audio-visual sound separation via hidden Markov models. In: Advances in Neural Information Processing Systems, pp. 1173\u20131180 (2002)"},{"key":"1918_CR9","unstructured":"Insta360.com: Insta360 360 Camera\u2014Insta360, the leader in 360 cameras. Retrieved November 19, 2019 from https:\/\/www.insta360.com\/"},{"key":"1918_CR10","doi-asserted-by":"crossref","unstructured":"Kim, D., Woo, S., Lee, J.Y., So\u00a0Kweon, I.: Deep video inpainting. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5792\u20135801 (2019)","DOI":"10.1109\/CVPR.2019.00594"},{"issue":"6","key":"1918_CR11","doi-asserted-by":"publisher","first-page":"1099","DOI":"10.1109\/TPAMI.2015.2477814","volume":"38","author":"S Korman","year":"2015","unstructured":"Korman, S., Avidan, S.: Coherency sensitive hashing. IEEE Trans. Pattern Anal. Mach. Intell. 38(6), 1099\u20131112 (2015)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"2","key":"1918_CR12","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1109\/MSP.2014.2369531","volume":"32","author":"K Kowalczyk","year":"2015","unstructured":"Kowalczyk, K., Thiergart, O., Taseska, M., Del Galdo, G., Pulkki, V., Habets, E.A.: Parametric spatial sound processing: a flexible and efficient solution to sound scene acquisition, modification, and reproduction. IEEE Signal Process. Mag. 32(2), 31\u201342 (2015)","journal-title":"IEEE Signal Process. Mag."},{"key":"1918_CR13","doi-asserted-by":"crossref","unstructured":"Le\u00a0Meur, O., Gautier, J., Guillemot, C.: Examplar-based inpainting based on local geometry. In: 2011 18th IEEE International Conference on Image Processing, pp. 3401\u20133404. IEEE (2011)","DOI":"10.1109\/ICIP.2011.6116441"},{"key":"1918_CR14","unstructured":"Morgado, P., Nvasconcelos, N., Langlois, T., Wang, O.: Self-supervised generation of spatial audio for 360 video. In: Advances in Neural Information Processing Systems, pp. 362\u2013372 (2018)"},{"key":"1918_CR15","doi-asserted-by":"crossref","unstructured":"Mroueh, Y., Marcheret, E., Goel, V.: Deep multimodal learning for audio-visual speech recognition. In: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 2130\u20132134. IEEE (2015)","DOI":"10.1109\/ICASSP.2015.7178347"},{"key":"1918_CR16","doi-asserted-by":"crossref","unstructured":"Nair, A.A., Reiter, A., Zheng, C., Nayar, S.: Audiovisual zooming: what you see is what you hear. In: Proceedings of the 27th ACM International Conference on Multimedia, pp. 1107\u20131118. ACM (2019)","DOI":"10.1145\/3343031.3351010"},{"key":"1918_CR17","doi-asserted-by":"crossref","unstructured":"Paredes, D., Rodriguez, P., Ragot, N.: Catadioptric omnidirectional image inpainting via a multi-scale approach and image unwrapping. In: 2013 IEEE International Symposium on Robotic and Sensors Environments (ROSE), pp. 67\u201372. IEEE (2013)","DOI":"10.1109\/ROSE.2013.6698420"},{"issue":"6","key":"1918_CR18","first-page":"503","volume":"55","author":"V Pulkki","year":"2007","unstructured":"Pulkki, V.: Spatial sound reproduction with directional audio coding. J. Audio Eng. Soc. 55(6), 503\u2013516 (2007)","journal-title":"J. Audio Eng. Soc."},{"key":"1918_CR19","unstructured":"Ricoh Company, Ltd.: 360-degree camera RICOH THETA. Retrieved November 19, 2019 from https:\/\/theta360.com\/"},{"issue":"3","key":"1918_CR20","doi-asserted-by":"publisher","first-page":"125","DOI":"10.1109\/MSP.2013.2296173","volume":"31","author":"B Rivet","year":"2014","unstructured":"Rivet, B., Wang, W., Naqvi, S.M., Chambers, J.A.: Audiovisual speech source separation: an overview of key methodologies. IEEE Signal Process. Mag. 31(3), 125\u2013134 (2014)","journal-title":"IEEE Signal Process. Mag."},{"key":"1918_CR21","doi-asserted-by":"crossref","unstructured":"Ruochen, W., Yuhong, Z., Wei, Z.: Acoustic zooming based on real-time metadata control. In: 2014 4th IEEE International Conference on Network Infrastructure and Digital Content, pp. 338\u2013342. IEEE (2014)","DOI":"10.1109\/ICNIDC.2014.7000321"},{"issue":"2","key":"1918_CR22","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1109\/53.665","volume":"5","author":"BD Van Veen","year":"1988","unstructured":"Van Veen, B.D., Buckley, K.M.: Beamforming: a versatile approach to spatial filtering. IEEE ASSP mag. 5(2), 4\u201324 (1988)","journal-title":"IEEE ASSP mag."},{"key":"1918_CR23","doi-asserted-by":"crossref","unstructured":"Upenik, E., Akyazi, P., Tuzmen, M., Ebrahimi, T.: Inpainting in omnidirectional images for privacy protection. In: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 2487\u20132491. IEEE (2019)","DOI":"10.1109\/ICASSP.2019.8683346"},{"issue":"9","key":"1918_CR24","first-page":"709","volume":"57","author":"J Vilkamo","year":"2009","unstructured":"Vilkamo, J., Lokki, T., Pulkki, V.: Directional audio coding: virtual microphone-based synthesis and subjective evaluation. J. Audio Eng. Soc. 57(9), 709\u2013724 (2009)","journal-title":"J. Audio Eng. Soc."},{"key":"1918_CR25","doi-asserted-by":"crossref","unstructured":"Wang, Q., Zhang, L., Bertinetto, L., Hu, W., Torr, P.H.: Fast online object tracking and segmentation: a unifying approach. In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2019)","DOI":"10.1109\/CVPR.2019.00142"},{"key":"1918_CR26","doi-asserted-by":"crossref","unstructured":"Xu, R., Li, X., Zhou, B., Loy, C.C.: Deep flow-guided video inpainting. In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2019)","DOI":"10.1109\/CVPR.2019.00384"},{"key":"1918_CR27","doi-asserted-by":"crossref","unstructured":"Yang, W., Qian, Y., K\u00e4m\u00e4r\u00e4inen, J.K., Cricri, F., Fan, L.: Object detection in equirectangular panorama. In: 2018 24th International Conference on Pattern Recognition (ICPR), pp. 2190\u20132195. IEEE (2018)","DOI":"10.1109\/ICPR.2018.8546070"},{"key":"1918_CR28","doi-asserted-by":"crossref","unstructured":"Yu, J., Lin, Z., Yang, J., Shen, X., Lu, X., Huang, T.S.: Generative image inpainting with contextual attention. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5505\u20135514 (2018)","DOI":"10.1109\/CVPR.2018.00577"},{"key":"1918_CR29","doi-asserted-by":"crossref","unstructured":"Zioulis, N., Karakottas, A., Zarpalas, D., Daras, P.: Omnidepth: Dense depth estimation for indoors spherical panoramas. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 448\u2013465 (2018)","DOI":"10.1007\/978-3-030-01231-1_28"}],"container-title":["The Visual Computer"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00371-020-01918-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00371-020-01918-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00371-020-01918-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,11]],"date-time":"2024-08-11T03:04:58Z","timestamp":1723345498000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00371-020-01918-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,7,31]]},"references-count":29,"journal-issue":{"issue":"10-12","published-print":{"date-parts":[[2020,10]]}},"alternative-id":["1918"],"URL":"https:\/\/doi.org\/10.1007\/s00371-020-01918-1","relation":{},"ISSN":["0178-2789","1432-2315"],"issn-type":[{"value":"0178-2789","type":"print"},{"value":"1432-2315","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,7,31]]},"assertion":[{"value":"31 July 2020","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Compliance with ethical standards"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}