{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,21]],"date-time":"2025-06-21T04:12:40Z","timestamp":1750479160839,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":37,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,10,15]],"date-time":"2019-10-15T00:00:00Z","timestamp":1571097600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,10,15]]},"DOI":"10.1145\/3343031.3351010","type":"proceedings-article","created":{"date-parts":[[2019,10,21]],"date-time":"2019-10-21T16:32:26Z","timestamp":1571675546000},"page":"1107-1118","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["Audiovisual Zooming"],"prefix":"10.1145","author":[{"given":"Arun Asokan","family":"Nair","sequence":"first","affiliation":[{"name":"Snap Research, Baltimore, MD, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Austin","family":"Reiter","sequence":"additional","affiliation":[{"name":"Snap Research, New York, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Changxi","family":"Zheng","sequence":"additional","affiliation":[{"name":"Snap Research, New York, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shree","family":"Nayar","sequence":"additional","affiliation":[{"name":"Snap Research, New York, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,10,15]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Joon Son Chung, and Andrew Zisserman","author":"Afouras Triantafyllos","year":"2018","unstructured":"Triantafyllos Afouras , Joon Son Chung, and Andrew Zisserman . 2018 . The Conversation : Deep Audio-Visual Speech Enhancement . arXiv preprint arXiv:1804.04121 (2018). Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. 2018. The Conversation: Deep Audio-Visual Speech Enhancement. arXiv preprint arXiv:1804.04121 (2018)."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/0022-460X(76)90552-6"},{"volume-title":"Microphone arrays: signal processing techniques and applications","author":"Brandstein Michael","key":"e_1_3_2_2_3_1","unstructured":"Michael Brandstein and Darren Ward . 2013. Microphone arrays: signal processing techniques and applications . Springer Science & Business Media . Michael Brandstein and Darren Ward. 2013. Microphone arrays: signal processing techniques and applications .Springer Science & Business Media."},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/PROC.1969.7278"},{"volume-title":"Audio Zoom for Smartphones Based On Multiple Adaptive Beamformers. In International Conference on Latent Variable Analysis and Signal Separation .","author":"Duong N.Q.K.","key":"e_1_3_2_2_5_1","unstructured":"N.Q.K. Duong , P. Berthet , S. Zabre , M. Kerdranvat , A. Ozerov , and L. Chevallier . 2017 . Audio Zoom for Smartphones Based On Multiple Adaptive Beamformers. In International Conference on Latent Variable Analysis and Signal Separation . N.Q.K. Duong, P. Berthet, S. Zabre, M. Kerdranvat, A. Ozerov, and L. Chevallier. 2017. Audio Zoom for Smartphones Based On Multiple Adaptive Beamformers. In International Conference on Latent Variable Analysis and Signal Separation ."},{"key":"e_1_3_2_2_6_1","volume-title":"Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation. In ACM Transactions on Graphics (SIGGRAPH) .","author":"Ephrat A.","year":"2018","unstructured":"A. Ephrat , I. Mosseri , O. Lang , T. Dekel , K. Wilson , A. Hassidim , W.T. Freeman , and M. Rubinstein . 2018 . Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation. In ACM Transactions on Graphics (SIGGRAPH) . A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W.T. Freeman, and M. Rubinstein. 2018. Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation. In ACM Transactions on Graphics (SIGGRAPH) ."},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2017.7965918"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2016.2647702"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAP.1982.1142739"},{"key":"e_1_3_2_2_10_1","first-page":"7","article-title":"Robust Adaptive Beamforming Based on Interference Covariance Matrix Reconstruction and Steering Vector Estimation","volume":"60","author":"Gu Y.","year":"2012","unstructured":"Y. Gu and A. Leshem . 2012 . Robust Adaptive Beamforming Based on Interference Covariance Matrix Reconstruction and Steering Vector Estimation . IEEE Transactions on Signal Processing , Vol. 60 , 7 (July 2012), 3881--3885. Y. Gu and A. Leshem. 2012. Robust Adaptive Beamforming Based on Interference Covariance Matrix Reconstruction and Steering Vector Estimation. IEEE Transactions on Signal Processing , Vol. 60, 7 (July 2012), 3881--3885.","journal-title":"IEEE Transactions on Signal Processing"},{"key":"e_1_3_2_2_11_1","unstructured":"John R Hershey and Michael Casey. 2002. Audio-visual sound separation via hidden Markov models. In Advances in Neural Information Processing Systems. 1173--1180.  John R Hershey and Michael Casey. 2002. Audio-visual sound separation via hidden Markov models. In Advances in Neural Information Processing Systems. 1173--1180."},{"volume-title":"Neural Network Based Spectral Mask Estimation For Acoustic Beamforming. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) .","author":"Heymann J.","key":"e_1_3_2_2_12_1","unstructured":"J. Heymann , L. Drude , and R. Haeb-Umbach . 2016 . Neural Network Based Spectral Mask Estimation For Acoustic Beamforming. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . J. Heymann, L. Drude, and R. Haeb-Umbach. 2016. Neural Network Based Spectral Mask Estimation For Acoustic Beamforming. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) ."},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAES.1983.309349"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2008-644"},{"key":"e_1_3_2_2_15_1","volume-title":"Sven Erik Nordholm, and Yee Hong Leung","author":"Lai Chiong Ching","year":"2017","unstructured":"Chiong Ching Lai , Sven Erik Nordholm, and Yee Hong Leung . 2017 . A Study Into the Design of Steerable Microphone Arrays .Springer. Chiong Ching Lai, Sven Erik Nordholm, and Yee Hong Leung. 2017. A Study Into the Design of Steerable Microphone Arrays .Springer."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201391"},{"key":"e_1_3_2_2_17_1","volume-title":"On robust Capon beamforming and diagonal loading","author":"Li Jian","year":"2003","unstructured":"Jian Li , Petre Stoica , and Zhisong Wang . 2003. On robust Capon beamforming and diagonal loading . IEEE transactions on signal processing , Vol. 51 , 7 ( 2003 ), 1702--1715. Jian Li, Petre Stoica, and Zhisong Wang. 2003. On robust Capon beamforming and diagonal loading. IEEE transactions on signal processing , Vol. 51, 7 (2003), 1702--1715."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/343374"},{"volume-title":"Berlin Beamforming Conference","author":"Ulf","key":"e_1_3_2_2_19_1","unstructured":"Ulf Michel et almbox. 2006. History of acoustic beamforming . In Berlin Beamforming Conference , Berlin, Germany, Nov. 21--22. Ulf Michel et almbox. 2006. History of acoustic beamforming. In Berlin Beamforming Conference, Berlin, Germany, Nov. 21--22."},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178347"},{"key":"e_1_3_2_2_21_1","unstructured":"Alan V Oppenheim. 1999. Discrete-time signal processing .Pearson Education India.  Alan V Oppenheim. 1999. Discrete-time signal processing .Pearson Education India."},{"key":"e_1_3_2_2_22_1","volume-title":"Efros","author":"Owens Andrew","year":"2018","unstructured":"Andrew Owens and Alexei A . Efros . 2018 . Audio-Visual Scene Analysis with Self-Supervised Multisensory Features. In ECCV . Springer , 639--658. Andrew Owens and Alexei A. Efros. 2018. Audio-Visual Scene Analysis with Self-Supervised Multisensory Features. In ECCV . Springer, 639--658."},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"crossref","unstructured":"V. Rabinovich and N. Alexandrov. 2013. Typical Array Geometries and Basic Beam Steering Methods. In Antenna Arrays and Automotive Applications. Springer Science  V. Rabinovich and N. Alexandrov. 2013. Typical Array Geometries and Basic Beam Steering Methods. In Antenna Arrays and Automotive Applications. Springer Science","DOI":"10.1007\/978-1-4614-1074-4"},{"key":"e_1_3_2_2_24_1","unstructured":"Business Media New York Chapter 2 23--54.  Business Media New York Chapter 2 23--54."},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2013.2296173"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2001.941023"},{"volume-title":"Proceedings of IC-NIDC .","author":"Ruochen W.","key":"e_1_3_2_2_27_1","unstructured":"W. Ruochen , Z. Yuhong , and Z. Wei . 2014. Acoustic Zooming Based on Real-Time Metadata Control . In Proceedings of IC-NIDC . W. Ruochen, Z. Yuhong, and Z. Wei. 2014. Acoustic Zooming Based on Real-Time Metadata Control. In Proceedings of IC-NIDC ."},{"key":"e_1_3_2_2_28_1","volume-title":"et almbox","author":"Stoica Petre","year":"2005","unstructured":"Petre Stoica , Randolph L Moses , et almbox . 2005 . Spectral analysis of signals. (2005). Petre Stoica, Randolph L Moses, et almbox. 2005. Spectral analysis of signals. (2005)."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2010.5495701"},{"volume-title":"Modal array signal processing: principles and applications of acoustic wavefield decomposition","author":"Teutsch Heinz","key":"e_1_3_2_2_30_1","unstructured":"Heinz Teutsch . 2007. Modal array signal processing: principles and applications of acoustic wavefield decomposition . Vol. 348 . Springer . Heinz Teutsch. 2007. Modal array signal processing: principles and applications of acoustic wavefield decomposition . Vol. 348. Springer."},{"key":"e_1_3_2_2_31_1","volume-title":"An Acoustical Zoom Based on Informed Spatial Filtering. In International Workshop on Acoustic Signal Enhancement (IWAENC) .","author":"Thiergart O.","year":"2014","unstructured":"O. Thiergart , K. Kowalczyk , and E.A.P. Habets . 2014 . An Acoustical Zoom Based on Informed Spatial Filtering. In International Workshop on Acoustic Signal Enhancement (IWAENC) . O. Thiergart, K. Kowalczyk, and E.A.P. Habets. 2014. An Acoustical Zoom Based on Informed Spatial Filtering. In International Workshop on Acoustic Signal Enhancement (IWAENC) ."},{"volume-title":"Detection, Estimation, and Modulation Theory","author":"Van Trees H.L.","key":"e_1_3_2_2_32_1","unstructured":"H.L. Van Trees . 2002. Detection, Estimation, and Modulation Theory , Part IV: Optimum Array Processing .Wiley, New York . H.L. Van Trees. 2002. Detection, Estimation, and Modulation Theory, Part IV: Optimum Array Processing .Wiley, New York."},{"key":"e_1_3_2_2_33_1","first-page":"2","article-title":"Beamforming: A versatile approach to spatial filtering. IEEE Acoust., Speech","volume":"5","author":"Van Veen B.D.","year":"1988","unstructured":"B.D. Van Veen and K.M. Buckley . 1988 . Beamforming: A versatile approach to spatial filtering. IEEE Acoust., Speech , Signal Process. Mag. , Vol. 5 , 2 (April 1988), 4--24. B.D. Van Veen and K.M. Buckley. 1988. Beamforming: A versatile approach to spatial filtering. IEEE Acoust., Speech, Signal Process. Mag. , Vol. 5, 2 (April 1988), 4--24.","journal-title":"Signal Process. Mag."},{"key":"e_1_3_2_2_34_1","first-page":"4","article-title":"Performance Measurement in Blind Audio Source Separation","volume":"14","author":"Vincent E.","year":"2006","unstructured":"E. Vincent , R. Gribonval , and C. Fevotte . 2006 . Performance Measurement in Blind Audio Source Separation . Transactions on Audio, Speech and Language Processing , Vol. 14 , 4 (July 2006), 1462--1469. E. Vincent, R. Gribonval, and C. Fevotte. 2006. Performance Measurement in Blind Audio Source Separation. Transactions on Audio, Speech and Language Processing , Vol. 14, 4 (July 2006), 1462--1469.","journal-title":"Transactions on Audio, Speech and Language Processing"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2007.898454"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/78.482097"},{"key":"e_1_3_2_2_37_1","volume-title":"The Sound of Pixels. In The European Conference on Computer Vision (ECCV) .","author":"Zhao Hang","year":"2018","unstructured":"Hang Zhao , Chuang Gan , Andrew Rouditchenko , Carl Vondrick , Josh McDermott , and Antonio Torralba . 2018 . The Sound of Pixels. In The European Conference on Computer Vision (ECCV) . Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. 2018. The Sound of Pixels. In The European Conference on Computer Vision (ECCV) ."}],"event":{"name":"MM '19: The 27th ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Nice France","acronym":"MM '19"},"container-title":["Proceedings of the 27th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3343031.3351010","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3343031.3351010","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:13:11Z","timestamp":1750201991000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3343031.3351010"}},"subtitle":["What You See Is What You Hear"],"short-title":[],"issued":{"date-parts":[[2019,10,15]]},"references-count":37,"alternative-id":["10.1145\/3343031.3351010","10.1145\/3343031"],"URL":"https:\/\/doi.org\/10.1145\/3343031.3351010","relation":{},"subject":[],"published":{"date-parts":[[2019,10,15]]},"assertion":[{"value":"2019-10-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}