{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T17:29:36Z","timestamp":1777656576307,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":54,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"STCSM","award":["No. 18DZ2270700, No. 21DZ1100100"],"award-info":[{"award-number":["No. 18DZ2270700, No. 21DZ1100100"]}]},{"name":"National Key R&D Program of China","award":["No. 2020YFB1406801"],"award-info":[{"award-number":["No. 2020YFB1406801"]}]},{"name":"State Key Laboratory of UHD Video and Audio Production and Presentation"},{"name":"111 plan","award":["No. BP0719010"],"award-info":[{"award-number":["No. BP0719010"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3548317","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:43:12Z","timestamp":1665416592000},"page":"3742-3753","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":31,"title":["Exploiting Transformation Invariance and Equivariance for Self-supervised Sound Localisation"],"prefix":"10.1145","author":[{"given":"Jinxiang","family":"Liu","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chen","family":"Ju","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weidi","family":"Xie","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University &amp; Shanghai AI Laboratory, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ya","family":"Zhang","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University &amp; Shanghai AI Laboratory, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58523-5_13"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_27"},{"key":"e_1_3_2_2_3_1","volume-title":"Soundnet: Learning sound representations from unlabeled video. In Advances in Neural Information Processing Systems.","author":"Aytar Yusuf","year":"2016","unstructured":"Yusuf Aytar , Carl Vondrick , and Antonio Torralba . 2016 . Soundnet: Learning sound representations from unlabeled video. In Advances in Neural Information Processing Systems. Yusuf Aytar, Carl Vondrick, and Antonio Torralba. 2016. Soundnet: Learning sound representations from unlabeled video. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-48881-3_56"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01659"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053174"},{"key":"e_1_3_2_2_8_1","volume-title":"International conference on machine learning. PMLR, 1597--1607","author":"Chen Ting","year":"2020","unstructured":"Ting Chen , Simon Kornblith , Mohammad Norouzi , and Geoffrey Hinton . 2020 a. A simple framework for contrastive learning of visual representations . In International conference on machine learning. PMLR, 1597--1607 . Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020a. A simple framework for contrastive learning of visual representations. In International conference on machine learning. PMLR, 1597--1607."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01549"},{"key":"e_1_3_2_2_10_1","volume-title":"Learning joint statistical models for audio-visual fusion and segregation. Advances in neural information processing systems","author":"John W","year":"2000","unstructured":"John W Fisher III, Trevor Darrell , William Freeman , and Paul Viola . 2000. Learning joint statistical models for audio-visual fusion and segregation. Advances in neural information processing systems , Vol. 13 ( 2000 ). John W Fisher III, Trevor Darrell, William Freeman, and Paul Viola. 2000. Learning joint statistical models for audio-visual fusion and segregation. Advances in neural information processing systems, Vol. 13 (2000)."},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01049"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01219-9_3"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00398"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01524"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01047"},{"key":"e_1_3_2_2_16_1","first-page":"21271","article-title":"Bootstrap your own latent-a new approach to self-supervised learning","volume":"33","author":"Grill Jean-Bastien","year":"2020","unstructured":"Jean-Bastien Grill , Florian Strub , Florent Altch\u00e9 , Corentin Tallec , Pierre Richemond , Elena Buchatskaya , Carl Doersch , Bernardo Avila Pires , Zhaohan Guo , Mohammad Gheshlaghi Azar , 2020 . Bootstrap your own latent-a new approach to self-supervised learning . Advances in Neural Information Processing Systems , Vol. 33 (2020), 21271 -- 21284 . Jean-Bastien Grill, Florian Strub, Florent Altch\u00e9, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. 2020. Bootstrap your own latent-a new approach to self-supervised learning. Advances in Neural Information Processing Systems, Vol. 33 (2020), 21271--21284.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_2_2_18_1","volume-title":"Audio vision: Using audio-visual synchrony to locate sounds. Advances in neural information processing systems","author":"Hershey John","year":"1999","unstructured":"John Hershey and Javier Movellan . 1999. Audio vision: Using audio-visual synchrony to locate sounds. Advances in neural information processing systems , Vol. 12 ( 1999 ). John Hershey and Javier Movellan. 1999. Audio vision: Using audio-visual synchrony to locate sounds. Advances in neural information processing systems, Vol. 12 (1999)."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24261-3_7"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00167"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00947"},{"key":"e_1_3_2_2_22_1","first-page":"10077","article-title":"Discriminative sounding objects localization via self-supervised audiovisual matching","volume":"33","author":"Hu Di","year":"2020","unstructured":"Di Hu , Rui Qian , Minyue Jiang , Xiao Tan , Shilei Wen , Errui Ding , Weiyao Lin , and Dejing Dou . 2020 . Discriminative sounding objects localization via self-supervised audiovisual matching . Advances in Neural Information Processing Systems , Vol. 33 (2020), 10077 -- 10087 . Di Hu, Rui Qian, Minyue Jiang, Xiao Tan, Shilei Wen, Errui Ding, Weiyao Lin, and Dejing Dou. 2020. Discriminative sounding objects localization via self-supervised audiovisual matching. Advances in Neural Information Processing Systems, Vol. 33 (2020), 10077--10087.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_23_1","volume-title":"Class-aware Sounding Objects Localization via Audiovisual Correspondence","author":"Hu Di","year":"2021","unstructured":"Di Hu , Yake Wei , Rui Qian , Weiyao Lin , Ruihua Song , and Ji-Rong Wen . 2021. Class-aware Sounding Objects Localization via Audiovisual Correspondence . IEEE Transactions on Pattern Analysis and Machine Intelligence ( 2021 ). Di Hu, Yake Wei, Rui Qian, Weiyao Lin, Ruihua Song, and Ji-Rong Wen. 2021. Class-aware Sounding Objects Localization via Audiovisual Correspondence. IEEE Transactions on Pattern Analysis and Machine Intelligence (2021)."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00559"},{"key":"e_1_3_2_2_25_1","unstructured":"Gregory Koch Richard Zemel Ruslan Salakhutdinov etal 2015. Siamese neural networks for one-shot image recognition. In ICML deep learning workshop Vol. 2. Lille 0.  Gregory Koch Richard Zemel Ruslan Salakhutdinov et al. 2015. Siamese neural networks for one-shot image recognition. In ICML deep learning workshop Vol. 2. Lille 0."},{"key":"e_1_3_2_2_26_1","volume-title":"Advances in Neural Information Processing Systems","volume":"31","author":"Korbar Bruno","year":"2018","unstructured":"Bruno Korbar , Du Tran , and Lorenzo Torresani . 2018 . Cooperative learning of audio and video models from self-supervised synchronization . Advances in Neural Information Processing Systems , Vol. 31 (2018). Bruno Korbar, Du Tran, and Lorenzo Torresani. 2018. Cooperative learning of audio and video models from self-supervised synchronization. Advances in Neural Information Processing Systems, Vol. 31 (2018)."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2016.59"},{"key":"e_1_3_2_2_28_1","volume-title":"International Conference on Learning Representations.","author":"Lee Jun-Tae","year":"2020","unstructured":"Jun-Tae Lee , Mihir Jain , Hyoungwoo Park , and Sungrack Yun . 2020 . Cross-attentional audio-visual fusion for weakly-supervised action localization . In International Conference on Learning Representations. Jun-Tae Lee, Mihir Jain, Hyoungwoo Park, and Sungrack Yun. 2020. Cross-attentional audio-visual fusion for weakly-supervised action localization. In International Conference on Learning Representations."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683226"},{"key":"e_1_3_2_2_30_1","volume-title":"Unsupervised sound localization via iterative contrastive learning. arXiv preprint arXiv:2104.00315","author":"Lin Yan-Bo","year":"2021","unstructured":"Yan-Bo Lin , Hung-Yu Tseng , Hsin-Ying Lee , Yen-Yu Lin , and Ming-Hsuan Yang . 2021. Unsupervised sound localization via iterative contrastive learning. arXiv preprint arXiv:2104.00315 ( 2021 ). Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin, and Ming-Hsuan Yang. 2021. Unsupervised sound localization via iterative contrastive learning. arXiv preprint arXiv:2104.00315 (2021)."},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3184558.3186942"},{"key":"e_1_3_2_2_32_1","volume-title":"Semi-supervised Keypoint Localization. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=yFJ67zTeI2","author":"Moskvyak Olga","year":"2021","unstructured":"Olga Moskvyak , Frederic Maire , Feras Dayoub , and Mahsa Baktashmotlagh . 2021 . Semi-supervised Keypoint Localization. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=yFJ67zTeI2 Olga Moskvyak, Frederic Maire, Feras Dayoub, and Mahsa Baktashmotlagh. 2021. Semi-supervised Keypoint Localization. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=yFJ67zTeI2"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58523-5_35"},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01231-1_39"},{"key":"e_1_3_2_2_35_1","volume-title":"Specaugment: A simple data augmentation method for automatic speech recognition.","author":"Park Daniel S","year":"2019","unstructured":"Daniel S Park , William Chan , Yu Zhang , Chung-Cheng Chiu , Barret Zoph , Ekin D Cubuk , and Quoc V Le . 2019 . Specaugment: A simple data augmentation method for automatic speech recognition. (2019), 2613--2617. Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le. 2019. Specaugment: A simple data augmentation method for automatic speech recognition. (2019), 2613--2617."},{"key":"e_1_3_2_2_36_1","volume-title":"Crafting Better Contrastive Views for Siamese Representation Learning. arXiv preprint arXiv:2202.03278","author":"Peng Xiangyu","year":"2022","unstructured":"Xiangyu Peng , Kai Wang , Zheng Zhu , and Yang You . 2022. Crafting Better Contrastive Views for Siamese Representation Learning. arXiv preprint arXiv:2202.03278 ( 2022 ). Xiangyu Peng, Kai Wang, Zheng Zhu, and Yang You. 2022. Crafting Better Contrastive Views for Siamese Representation Learning. arXiv preprint arXiv:2202.03278 (2022)."},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58565-5_18"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00458"},{"key":"e_1_3_2_2_39_1","volume-title":"Learning to localize sound sources in visual scenes: Analysis and applications","author":"Senocak Arda","year":"2019","unstructured":"Arda Senocak , Tae-Hyun Oh , Junsik Kim , Ming-Hsuan Yang , and In So Kweon . 2019. Learning to localize sound sources in visual scenes: Analysis and applications . IEEE transactions on pattern analysis and machine intelligence, Vol. 43 , 5 ( 2019 ), 1605--1619. Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon. 2019. Learning to localize sound sources in visual scenes: Analysis and applications. IEEE transactions on pattern analysis and machine intelligence, Vol. 43, 5 (2019), 1605--1619."},{"key":"e_1_3_2_2_40_1","volume-title":"Learning Sound Localization Better From Semantically Similar Samples. In ICASSP 2022--2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE.","author":"Senocak Arda","year":"2022","unstructured":"Arda Senocak , Hyeonggon Ryu , Junsik Kim , and In So Kweon . 2022 . Learning Sound Localization Better From Semantically Similar Samples. In ICASSP 2022--2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE. Arda Senocak, Hyeonggon Ryu, Junsik Kim, and In So Kweon. 2022. Learning Sound Localization Better From Semantically Similar Samples. In ICASSP 2022--2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE."},{"key":"e_1_3_2_2_41_1","volume-title":"Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR).","author":"Song Zengjie","year":"2022","unstructured":"Zengjie Song , Yuxi Wang , Junsong Fan , Tieniu Tan , and Zhaoxiang Zhang . 2022 . Self-Supervised Predictive Learning: A Negative-Free Method for Sound Source Localization in Visual Scenes . In Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR). Zengjie Song, Yuxi Wang, Junsong Fan, Tieniu Tan, and Zhaoxiang Zhang. 2022. Self-Supervised Predictive Learning: A Negative-Free Method for Sound Source Localization in Visual Scenes. In Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.220"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00646"},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.348"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58580-8_26"},{"key":"e_1_3_2_2_46_1","volume-title":"Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds. In 9th International Conference on Learning Representations, ICLR 2021","author":"Tzinis Efthymios","year":"2021","unstructured":"Efthymios Tzinis , Scott Wisdom , Aren Jansen , Shawn Hershey , Tal Remez , Dan Ellis , and John R. Hershey . 2021 . Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds. In 9th International Conference on Learning Representations, ICLR 2021 , Virtual Event, Austria, May 3--7 , 2021 . OpenReview.net. https:\/\/openreview.net\/forum?id=MDsQkFP1Aw Efthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey, Tal Remez, Dan Ellis, and John R. Hershey. 2021. Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3--7, 2021. OpenReview.net. https:\/\/openreview.net\/forum?id=MDsQkFP1Aw"},{"key":"e_1_3_2_2_47_1","volume-title":"Representation Learning with Contrastive Predictive Coding. CoRR","author":"den Oord Aaron Van","year":"2018","unstructured":"Aaron Van den Oord , Yazhe Li , and Oriol Vinyals . 2018. Representation Learning with Contrastive Predictive Coding. CoRR , Vol. abs\/ 1807 .03748 ( 2018 ). showeprint[arXiv]1807.03748 http:\/\/arxiv.org\/abs\/1807.03748 Aaron Van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation Learning with Contrastive Predictive Coding. CoRR, Vol. abs\/1807.03748 (2018). showeprint[arXiv]1807.03748 http:\/\/arxiv.org\/abs\/1807.03748"},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01271"},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01229"},{"key":"e_1_3_2_2_50_1","volume-title":"Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer.","author":"Xiao Fanyi","year":"2020","unstructured":"Fanyi Xiao , Yong Jae Lee , Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer. 2020 . Audiovisual slowfast networks for video recognition. arXiv preprint arXiv:2001.08740 (2020). Fanyi Xiao, Yong Jae Lee, Kristen Grauman, Jitendra Malik, and Christoph Feichtenhofer. 2020. Audiovisual slowfast networks for video recognition. arXiv preprint arXiv:2001.08740 (2020)."},{"key":"e_1_3_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00097"},{"key":"e_1_3_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00182"},{"key":"e_1_3_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_35"},{"key":"e_1_3_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.319"}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548317","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3548317","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:43Z","timestamp":1750186843000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548317"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":54,"alternative-id":["10.1145\/3503161.3548317","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3548317","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}