{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,1]],"date-time":"2026-08-01T02:38:16Z","timestamp":1785551896301,"version":"3.56.0"},"reference-count":77,"publisher":"Association for Computing Machinery (ACM)","issue":"6","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,6,30]]},"abstract":"<jats:p>\n            Audio-Visual Semantic Segmentation (AVSS) plays a crucial role in pixel-level multi-modal perception for real-world applications such as robotic navigation and autonomous driving. Existing methods typically rely on global spatio-temporal modules to fuse audio and visual representations, which aids in generating pixel-level semantic masks. However, these approaches often overlook the importance of local spatio-temporal context in understanding semantics, leading to suboptimal performance. This limitation makes it difficult for models to accurately distinguish sound-emitting objects from irrelevant background noise, resulting in erroneous segmentation across the spatio-temporal dimension. To address this issue, we propose the\n            <jats:italic toggle=\"yes\">ALOHA<\/jats:italic>\n            framework, which\n            <jats:italic toggle=\"yes\">A<\/jats:italic>\n            dapts\n            <jats:italic toggle=\"yes\">LO<\/jats:italic>\n            cal spatio-temporal context to en\n            <jats:italic toggle=\"yes\">HA<\/jats:italic>\n            nce AVSS. The framework introduces two key components designed to leverage and enhance local spatio-temporal context information: the LOHA adapter and the Selective Context Enhancement (SCE) module. Specifically, the LOHA adapter adaptively captures essential modality information across spatio-temporal dimensions, while implicitly learning fine-grained local context through the local attention mechanism. Furthermore, the SCE module selectively enhances the local context related to the semantics, thereby facilitating the distinction between the sounding object and irrelevant background and improving segmentation accuracy. Moreover, to better adapt to embodied AI systems, our framework utilizes a parameter-shared encoder and applies the adapters in a staged manner. This design significantly reduces the number of trainable parameters, making it more parameter-efficient. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance on the AVSBench-Semantic benchmark dataset and shows competitive results on the AVSBench-Object benchmark, while exhibiting broad adaptability across different visual backbone networks.\n          <\/jats:p>","DOI":"10.1145\/3735975","type":"journal-article","created":{"date-parts":[[2025,5,22]],"date-time":"2025-05-22T17:19:21Z","timestamp":1747934361000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["ALOHA: Adapting Local Spatio-Temporal Context to Enhance the Audio-Visual Semantic Segmentation"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-9401-4432","authenticated-orcid":false,"given":"Yang-Hao","family":"Zhou","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0320-7520","authenticated-orcid":false,"given":"Heyan","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4546-3928","authenticated-orcid":false,"given":"Cunhan","family":"Guo","sequence":"additional","affiliation":[{"name":"School of Emergency Management Science and Engineering, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9567-159X","authenticated-orcid":false,"given":"Rong-Cheng","family":"Tu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9171-9234","authenticated-orcid":false,"given":"Zeyu","family":"Xiao","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8102-5346","authenticated-orcid":false,"given":"Bo","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6795-2311","authenticated-orcid":false,"given":"Xian-Ling","family":"Mao","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,7,8]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neuroimage.2010.05.002"},{"key":"e_1_3_1_3_2","volume-title":"The Cognitive Neuroscience of Mind: A Tribute to Michael S. Gazzaniga","author":"Gazzaniga Michael S.","year":"2010","unstructured":"Michael S. Gazzaniga. 2010. The Cognitive Neuroscience of Mind: A Tribute to Michael S. Gazzaniga. MIT Press."},{"issue":"3","key":"e_1_3_1_4_2","doi-asserted-by":"crossref","first-page":"241","DOI":"10.3758\/BF03203206","article-title":"Eye movements in auditory space perception","volume":"17","author":"Jones Bill","year":"1975","unstructured":"Bill Jones and Boris Kabanoff. 1975. Eye movements in auditory space perception. Perception & Psychophysics 17, 3 (1975), 241\u2013245.","journal-title":"Perception & Psychophysics"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01018"},{"key":"e_1_3_1_6_2","unstructured":"Arda Senocak Hyeonggon Ryu Junsik Kim Tae-Hyun Oh Hanspeter Pfister and Joon Son Chung. 2024. Aligning sight and sound: Advanced sound source localization through audio-visual alignment. arXiv:2407.13676. Retrieved from https:\/\/arxiv.org\/abs\/2407.13676"},{"key":"e_1_3_1_7_2","unstructured":"Jinxing Zhou Xuyang Shen Jianyuan Wang Jiayi Zhang Weixuan Sun Jing Zhang Stan Birchfield Dan Guo Lingpeng Kong Meng Wang et al. 2023. Audio-visual segmentation with semantics. arXiv:2301.13190. Retrieved from https:\/\/arxiv.org\/abs\/2301.13190"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2024.124885"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00228"},{"key":"e_1_3_1_10_2","article-title":"Cross-modal prompts: Adapting large pre-trained models for audio-visual downstream tasks","volume":"36","author":"Duan Haoyi","year":"2024","unstructured":"Haoyi Duan, Yan Xia, Zhou Mingze, Li Tang, Jieming Zhu, and Zhou Zhao. 2024. Cross-modal prompts: Adapting large pre-trained models for audio-visual downstream tasks. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 36.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_11_2","first-page":"954","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Mao Yuxin","year":"2023","unstructured":"Yuxin Mao, Jing Zhang, Mochu Xiang, Yiran Zhong, and Yuchao Dai. 2023. Multimodal variational auto-encoder based audio-visual segmentation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 954\u2013965."},{"key":"e_1_3_1_12_2","unstructured":"Shengyi Gao Zhe Chen Guo Chen Wenhai Wang and Tong Lu. 2023. AVSegFormer: Audio-visual segmentation with transformer. arXiv:2307.01146. Retrieved from https:\/\/arxiv.org\/abs\/2307.01146"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611724"},{"key":"e_1_3_1_14_2","unstructured":"Yuxuan Li Xiang Li Yimain Dai Qibin Hou Li Liu Yongxiang Liu Ming-Ming Cheng and Jian Yang. 2024. LSKNet: A foundation lightweight backbone for remote sensing. arXiv:2403.11735. Retrieved from https:\/\/arxiv.org\/abs\/2403.11735"},{"key":"e_1_3_1_15_2","unstructured":"Sihan Chen Handong Li Qunbo Wang Zijia Zhao Mingzhen Sun Xinxin Zhu and Jing Liu. 2023. VAST: A vision-audio-subtitle-text omni-modality foundation model and dataset. arXiv:2305.18500. Retrieved from https:\/\/arxiv.org\/abs\/2305.18500"},{"key":"e_1_3_1_16_2","volume-title":"Proceedings of the 11th International Conference on Learning Representations","author":"Gong Yuan","year":"2022","unstructured":"Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R. Glass. 2022. Contrastive audio-visual masked autoencoder. In Proceedings of the 11th International Conference on Learning Representations."},{"key":"e_1_3_1_17_2","first-page":"976","volume-title":"Proceedings of the 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201922)","author":"Guzhov Andrey","year":"2022","unstructured":"Andrey Guzhov, Federico Raue, J\u00f6rn Hees, and Andreas Dengel. 2022. AudioCLIP: Extending clip to image, text and audio. In Proceedings of the 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP \u201922). IEEE, 976\u2013980."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01526"},{"key":"e_1_3_1_19_2","first-page":"1","volume-title":"ACM Transactions on Multimedia Computing, Communications and Applications","volume":"20","author":"Hou Wenxuan","year":"2023","unstructured":"Wenxuan Hou, Guangyao Li, Yapeng Tian, and Di Hu. 2023. Towards long form audio-visual video understanding. ACM Transactions on Multimedia Computing, Communications and Applications 20, 9 (2023), 1\u201326."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01216-8_16"},{"key":"e_1_3_1_21_2","first-page":"436","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part III","author":"Tian Yapeng","year":"2020","unstructured":"Yapeng Tian, Dingzeyu Li, and Chenliang Xu. 2020. Unified multisensory perception: Weakly-supervised audio-visual video parsing. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part III. Springer, 436\u2013454."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00138"},{"key":"e_1_3_1_23_2","first-page":"1415","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Geng Shijie","year":"2021","unstructured":"Shijie Geng, Peng Gao, Moitreya Chatterjee, Chiori Hori, Jonathan Le Roux, Yongfeng Zhang, Hongsheng Li, and Anoop Cherian. 2021. Dynamic graph representation learning for video dialog via multi-modal shuffled transformers. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35, 1415\u20131423."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01852"},{"key":"e_1_3_1_25_2","first-page":"2031","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Yun Heeseung","year":"2021","unstructured":"Heeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee, and Gunhee Kim. 2021. Pano-AVQA: Grounded audio-visual question answering on 360\u00b0 videos. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 2031\u20132041."},{"key":"e_1_3_1_26_2","first-page":"386","volume-title":"Proceedings of the 17th European Conference on Computer Vision (ECCV \u201922), Part XXXVII","author":"Zhou Jinxing","year":"2022","unstructured":"Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong. 2022. Audio\u2013visual segmentation. In Proceedings of the 17th European Conference on Computer Vision (ECCV \u201922), Part XXXVII. Springer, 386\u2013403."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00060"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-022-0274-8"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00135"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00695"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00493"},{"key":"e_1_3_1_33_2","unstructured":"Jinxiang Liu Chen Ju Chaofan Ma Yanfeng Wang Yu Wang and Ya Zhang. 2023. Audio-aware query-enhanced transformer for audio-visual segmentation. arXiv:2307.13236. Retrieved from https:\/\/arxiv.org\/abs\/2307.13236"},{"key":"e_1_3_1_34_2","doi-asserted-by":"crossref","unstructured":"Shaofei Huang Yuqing Han Li Hongji Wang Jiao Zhu Jizhong Han Dai Wenge Rong and Si Liu. 2023. Discovering sounding objects by audio queries for audio visual segmentation. arXiv:2309.09501. Retrieved from https:\/\/arxiv.org\/abs\/2309.09501","DOI":"10.24963\/ijcai.2023\/97"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2024.3405622"},{"key":"e_1_3_1_36_2","first-page":"5604","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Liu Jinxiang","year":"2024","unstructured":"Jinxiang Liu, Yu Wang, Chen Ju, Chaofan Ma, Ya Zhang, and Weidi Xie. 2024. Annotation-free audio-visual segmentation. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 5604\u20135614."},{"key":"e_1_3_1_37_2","unstructured":"Shentong Mo and Yapeng Tian. 2023. AV-SAM: Segment anything model meets audio-visual localization and segmentation. arXiv:2305.01836. Retrieved from https:\/\/arxiv.org\/abs\/2305.01836"},{"key":"e_1_3_1_38_2","unstructured":"Yuxin Mao Jing Zhang Mochu Xiang Yunqiu Lv Yiran Zhong and Yuchao Dai. 2023. Contrastive conditional latent diffusion for audio-visual segmentation. arXiv:2307.16579. Retrieved from https:\/\/arxiv.org\/abs\/2307.16579"},{"key":"e_1_3_1_39_2","first-page":"26328","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu Jinxiang","year":"2024","unstructured":"Jinxiang Liu, Yikun Liu, Fei Zhang, Chen Ju, Ya Zhang, and Yanfeng Wang. 2024. Audio-visual segmentation via unlabeled frame exploitation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 26328\u201326339."},{"key":"e_1_3_1_40_2","first-page":"1950","article-title":"Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning","volume":"35","author":"Liu Haokun","year":"2022","unstructured":"Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A. Raffel. 2022. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 35, 1950\u20131965.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_41_2","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2021. LoRA: Low-rank adaptation of large language models. arXiv:2106.09685. Retrieved from https:\/\/arxiv.org\/abs\/2106.09685"},{"key":"e_1_3_1_42_2","first-page":"23826","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang Lingxiao","year":"2024","unstructured":"Lingxiao Yang, Ru-Yuan Zhang, Yanchen Wang, and Xiaohua Xie. 2024. MMA: Multi-modal adapter for vision-language models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 23826\u201323837."},{"key":"e_1_3_1_43_2","unstructured":"Renrui Zhang Jiaming Han Chris Liu Peng Gao Aojun Zhou Xiangfei Hu Shilin Yan Pan Lu Hongsheng Li and Yu Qiao. 2023. LLaMA-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv:2303.16199. Retrieved from https:\/\/arxiv.org\/abs\/2303.16199"},{"key":"e_1_3_1_44_2","unstructured":"Peng Gao Jiaming Han Renrui Zhang Ziyi Lin Shijie Geng Aojun Zhou Wei Zhang Pan Lu Conghui He Xiangyu Yue et al. 2023. LLaMA-adapter V2: Parameter-efficient visual instruction model. arXiv:2304.15010. Retrieved from https:\/\/arxiv.org\/abs\/2304.15010"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-023-01891-x"},{"key":"e_1_3_1_46_2","unstructured":"Renrui Zhang Rongyao Fang Wei Zhang Peng Gao Kunchang Li Jifeng Dai Yu Qiao and Hongsheng Li. 2021. Tip-adapter: Training-free clip-adapter for better vision-language modeling. arXiv:2111.03930. Retrieved from https:\/\/arxiv.org\/abs\/2111.03930"},{"key":"e_1_3_1_47_2","unstructured":"Zhaoxu Li Zitong Yu Nithish Muthuchamy Selvaraj Xiaobao Guo Bingquan Shen Adams Wai-Kin Kong and Alex Kot. 2023. Flexible-modal deception detection with audio-visual adapter. arXiv:2302.05727. Retrieved from https:\/\/arxiv.org\/abs\/2302.05727"},{"key":"e_1_3_1_48_2","first-page":"14200","article-title":"Attention bottlenecks for multimodal fusion","volume":"34","author":"Nagrani Arsha","year":"2021","unstructured":"Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun. 2021. Attention bottlenecks for multimodal fusion. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 34, 14200\u201314213.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3262578"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3613440"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612293"},{"key":"e_1_3_1_52_2","volume-title":"Proceedings of the ACM Multimedia 2024","author":"Du Yang","year":"2024","unstructured":"Yang Du, Yuqi Liu, and Qin Jin. 2024. Reversed in time: A novel temporal-emphasized benchmark for cross-modal video-text retrieval. In Proceedings of the ACM Multimedia 2024."},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612096"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3393452"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3377475"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3648368"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3656046"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01540"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3558520"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00656"},{"key":"e_1_3_1_61_2","first-page":"10840","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu Chenchen","year":"2020","unstructured":"Chenchen Liu, Yang Jin, Kehan Xu, Guoqiang Gong, and Yadong Mu. 2020. Beyond short-term snippet: Video relation detection with spatio-temporal global context. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 10840\u201310849."},{"key":"e_1_3_1_62_2","first-page":"10810","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Mun Jonghwan","year":"2020","unstructured":"Jonghwan Mun, Minsu Cho, and Bohyung Han. 2020. Local-global video-text interactions for temporal grounding. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 10810\u201310819."},{"key":"e_1_3_1_63_2","unstructured":"Rong-Cheng Tu Yatai Ji Jie Jiang Weijie Kong Chengfei Cai Wenzhe Zhao Hongfa Wang Yujiu Yang and Wei Liu. 2023. Global and local semantic completion learning for vision-language pre-training. arXiv:2306.07096. Retrieved from https:\/\/arxiv.org\/abs\/2306.07096"},{"key":"e_1_3_1_64_2","first-page":"616","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics","author":"Li Jingye","year":"2020","unstructured":"Jingye Li, Hao Fei, and Donghong Ji. 2020. Modeling local contexts for joint dialogue act recognition and sentiment classification with bi-channel dynamic convolutions. In Proceedings of the 28th International Conference on Computational Linguistics, 616\u2013626."},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612053"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3318967"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.73"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01055"},{"key":"e_1_3_1_69_2","unstructured":"Weihao Yu Chenyang Si Pan Zhou Mi Luo Yichen Zhou Jiashi Feng Shuicheng Yan and Xinchao Wang. 2022. Metaformer baselines for vision. arXiv:2210.13452. Retrieved from https:\/\/arxiv.org\/abs\/2210.13452"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00656"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"e_1_3_1_72_2","first-page":"2491","article-title":"Associating objects with transformers for video object segmentation","volume":"34","author":"Yang Zongxin","year":"2021","unstructured":"Zongxin Yang, Yunchao Wei, and Yi Yang. 2021. Associating objects with transformers for video object segmentation. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 34, 2491\u20132502.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01659"},{"key":"e_1_3_1_74_2","first-page":"292","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part XX","author":"Qian Rui","year":"2020","unstructured":"Rui Qian, Di Hu, Heinrich Dinkel, Mengyue Wu, Ning Xu, and Weiyao Lin. 2020. Multiple sound sources localization from coarse to fine. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920), Part XX. Springer, 292\u2013308."},{"key":"e_1_3_1_75_2","doi-asserted-by":"crossref","unstructured":"Sabarinath Mahadevan Ali Athar Aljo\u0161a O\u0161ep Sebastian Hennen Laura Leal-Taix\u00e9 and Bastian Leibe. 2020. Making a case for 3D convolutions for object segmentation in videos. arXiv:2008.11516. Retrieved from https:\/\/arxiv.org\/abs\/2008.11516","DOI":"10.5244\/C.34.65"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00585"},{"key":"e_1_3_1_77_2","first-page":"15448","article-title":"Learning generative vision transformer with energy-based latent space for saliency prediction","volume":"34","author":"Zhang Jing","year":"2021","unstructured":"Jing Zhang, Jianwen Xie, Nick Barnes, and Ping Li. 2021. Learning generative vision transformer with energy-based latent space for saliency prediction. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 34, 15448\u201315463.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_78_2","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556. Retrieved from https:\/\/arxiv.org\/abs\/1409.1556"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3735975","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,8]],"date-time":"2025-07-08T17:07:43Z","timestamp":1751994463000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3735975"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,30]]},"references-count":77,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,6,30]]}},"alternative-id":["10.1145\/3735975"],"URL":"https:\/\/doi.org\/10.1145\/3735975","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,30]]},"assertion":[{"value":"2024-11-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}