{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T15:10:21Z","timestamp":1784905821615,"version":"3.55.0"},"reference-count":82,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2025,4,7]],"date-time":"2025-04-07T00:00:00Z","timestamp":1743984000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62377011, 62302337, 62402098"],"award-info":[{"award-number":["62377011, 62302337, 62402098"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Foundation of Key Laboratory of Artificial Intelligence, Ministry of Education"},{"name":"Shanghai Super Doctoral Incentive Program"},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["2232024D-25"],"award-info":[{"award-number":["2232024D-25"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,4,30]]},"abstract":"<jats:p>With the remarkable advancement of deep learning techniques and the wide availability of large-scale datasets, the performance of audio-visual saliency prediction has been drastically improved. Actually, audio-visual saliency prediction is still at an early exploration stage due to the spatial-temporal signal complexity and dynamic continuity of video content. To our knowledge, most existing audio-visual saliency prediction approaches usually represent videos as 3D grid of RGB values using discrete convolutional neural networks (CNNs), which inevitably incurs video content-agnostic and ignores the dynamic continuity issues. This article proposes a novel parametric audio-visual saliency (PAVS) model with implicit neural representation (INR) to address the aforementioned problems. Specifically, by using the proposed parametric neural network, we can effectively encode the space-time coordinates of video frames into corresponding saliency values, which can significantly enhance the compact feature representation ability. Meanwhile, a parametric feature fusion method is developed to achieve intrinsic interactions between audio and visual information streams, which can adaptively fuse audio and visual features to obtain competitive performance. Notably, without resorting to any specific audio-visual feature fusion strategy, the proposed PAVS model outperforms other state-of-the-art saliency methods by a large margin.<\/jats:p>","DOI":"10.1145\/3698881","type":"journal-article","created":{"date-parts":[[2024,10,11]],"date-time":"2024-10-11T16:00:51Z","timestamp":1728662451000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Audio-Visual Saliency Prediction Model with Implicit Neural Representation"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9843-5269","authenticated-orcid":false,"given":"Nana","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Donghua University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-9393-8975","authenticated-orcid":false,"given":"Min","family":"Xiong","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Donghua University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0329-6321","authenticated-orcid":false,"given":"Dandan","family":"Zhu","sequence":"additional","affiliation":[{"name":"Institute of AI Education, East China Normal University, Shanghai, China and Key Laboratory of Artificial Intelligence, Ministry of Education, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5773-5089","authenticated-orcid":false,"given":"Kun","family":"Zhu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Embedded System and Service Computing, Ministry of Education, Tongji University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8165-9322","authenticated-orcid":false,"given":"Guangtao","family":"Zhai","sequence":"additional","affiliation":[{"name":"Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4029-3322","authenticated-orcid":false,"given":"Xiaokang","family":"Yang","sequence":"additional","affiliation":[{"name":"MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,4,7]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01405"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01519-y"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2815601"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/2981562.2981593"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00574"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3203421"},{"key":"e_1_3_1_8_2","first-page":"21557","volume-title":"Proceedings of the Advances in Neural Information Processing SystemsCurran Associates, Inc","volume":"34","author":"Chen Hao","year":"2021","unstructured":"Hao Chen, Bo He, Hanyu Wang, Yixuan Ren, Ser Nam Lim, and Abhinav Shrivastava. 2021. NeRV: Neural representations for videos. In Proceedings of the Advances in Neural Information Processing Systems. M. Ranzato, A. Beygelzimer, Y. Dauphin, P. S. Liang, and J. Wortman Vaughan (Eds.), Vol. 34. Curran Associates, Inc., 21557\u201321568. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2021\/file\/b44182379bf9fae976e6ae5996e13cd8-Paper.pdf"},{"key":"e_1_3_1_9_2","first-page":"1","volume-title":"Proceedings of the IEEE 4th International Conference on Multimedia Big Data (BigMM \u201918)","author":"Chen Yangyu","year":"2018","unstructured":"Yangyu Chen, W. Zhang, Shuhui Wang, L. Li, and Qingming Huang. 2018. Saliency-based spatiotemporal attention for video captioning. In Proceedings of the IEEE 4th International Conference on Multimedia Big Data (BigMM \u201918), 1\u20138. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:53045827"},{"key":"e_1_3_1_10_2","first-page":"1","volume-title":"Proceedings of the 14th International Workshop on Image Analysis for Multimedia Interactive Services (WIAMIS \u201913)","author":"Coutrot Antoine","year":"2013","unstructured":"Antoine Coutrot and Nathalie Guyader. 2013. Toward the introduction of auditory information in dynamic visual attention models. In Proceedings of the 14th International Workshop on Image Analysis for Multimedia Interactive Services (WIAMIS \u201913), 1\u20134."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1167\/14.8.5"},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","first-page":"291","DOI":"10.1007\/978-1-4939-3435-5_16","article-title":"Multimodal saliency models for videos","volume":"10","author":"Coutrot Antoine","year":"2016","unstructured":"Antoine Coutrot and Nathalie Guyader. 2016. Multimodal saliency models for videos. From Human Attention to Computational Attention: A Multidisciplinary Approach 10 (2016), 291\u2013304.","journal-title":"From Human Attention to Computational Attention: A Multidisciplinary Approach"},{"key":"e_1_3_1_13_2","first-page":"6970","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems","volume":"30","author":"Dabkowski Piotr","year":"2017","unstructured":"Piotr Dabkowski and Yarin Gal. 2017. Real time image saliency for black box classifiers. In Proceedings of the 31st International Conference on Neural Information Processing Systems, Vol. 30, 6970\u20136979."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/QoMEX.2017.7965634"},{"key":"e_1_3_1_15_2","first-page":"419","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Droste Richard","year":"2020","unstructured":"Richard Droste, Jianbo Jiao, and J Alison Noble. 2020. Unified image and video saliency modeling. In Proceedings of the European Conference on Computer Vision. Springer, 419\u2013435."},{"key":"e_1_3_1_16_2","first-page":"2989","volume-title":"Proceedings of the 25th International Conference on Artificial Intelligence and Statistics","volume":"151","author":"Dupont Emilien","year":"2022","unstructured":"Emilien Dupont, Yee Whye Teh, and Arnaud Doucet. 2022. Generative models as distributions of functions. In Proceedings of the 25th International Conference on Artificial Intelligence and Statistics, Vol. 151, 2989\u20133015."},{"issue":"5","key":"e_1_3_1_17_2","doi-asserted-by":"crossref","first-page":"2305","DOI":"10.1109\/TIP.2018.2885229","article-title":"Deep3DSaliency: Deep stereoscopic video saliency detection model by 3D convolutional networks","volume":"28","author":"Fang Yuming","year":"2018","unstructured":"Yuming Fang, Guanqun Ding, Jia Li, and Zhijun Fang. 2018. Deep3DSaliency: Deep stereoscopic video saliency detection model by 3D convolutional networks. IEEE Transactions on Image Processing 28, 5 (2018), 2305\u20132318.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10584-0_33"},{"key":"e_1_3_1_19_2","first-page":"545","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"19","author":"Harel Jonathan","year":"2007","unstructured":"Jonathan Harel, Christof Koch, and Pietro Perona. 2007. Graph-based visual saliency. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 19, 545\u2013552."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2007.383267"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.153"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3069812"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.38"},{"key":"e_1_3_1_24_2","first-page":"805","volume-title":"Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS \u201923)","author":"Iacono Massimiliano","year":"2019","unstructured":"Massimiliano Iacono, Giulia D\u2019Angelo, Arren J. Glover, Vadim Tikhanoff, Ernst Niebur, and Chiara Bartolozzi. 2019. Proto-object based saliency for event-driven cameras. In Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS \u201923), 805\u2013812. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:210971188"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2004.834657"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/34.730558"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS51168.2021.9635989"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.620"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2020.103887"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_37"},{"key":"e_1_3_1_31_2","first-page":"2106","volume-title":"Proceedings of the IEEE 12th International Conference on Computer Vision.","author":"Judd Tilke","year":"2009","unstructured":"Tilke Judd, Krista Ehinger, Fr\u00e9do Durand, and Antonio Torralba. 2009. Learning to predict where humans look. In Proceedings of the IEEE 12th International Conference on Computer Vision. IEEE, 2106\u20132113."},{"issue":"4","key":"e_1_3_1_32_2","first-page":"219","article-title":"Shifts in selective visual attention: Towards the underlying neural circuitry","volume":"4","author":"Koch Christof","year":"1985","unstructured":"Christof Koch and Shimon Ullman. 1985. Shifts in selective visual attention: Towards the underlying neural circuitry. Human Neurobiology 4, 4 (1985), 219\u2013227.","journal-title":"Human Neurobiology"},{"key":"e_1_3_1_33_2","first-page":"5742","volume-title":"Proceedings of International Conference on Machine Learning.","author":"Kosiorek Adam R.","year":"2021","unstructured":"Adam R. Kosiorek, Heiko Strathmann, Daniel Zoran, Pol Moreno, Rosalia Schneider, Sona Mokr\u00e1, and Danilo Jimenez Rezende. 2021. NeRF-VAE: A geometry aware 3D scene generative model. In Proceedings of International Conference on Machine Learning. PMLR, 5742\u20135752."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.image.2015.08.004"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","unstructured":"Matthias K\u00fcmmerer Lucas Theis and Matthias Bethge. 2014. Deep gaze I: Boosting saliency prediction with feature maps trained on imagenet. arXiv:1411.1045. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.1411.1045","DOI":"10.48550\/arXiv.1411.1045"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2936112"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00544"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2020.3023080"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.126432"},{"key":"e_1_3_1_40_2","volume-title":"Proceedings of the 30th British Machine Vision Conference 2019 (BMVC \u201919)","volume":"182","author":"Linardos Panagiotis","year":"2019","unstructured":"Panagiotis Linardos, Eva Mohedano, Juan Jos\u00e9 Nieto, Noel E. O\u2019Connor, Xavier Gir\u00f3-i-Nieto, and Kevin McGuinness. 2019. Simple vs complex temporal recurrences for video saliency prediction. In Proceedings of the 30th British Machine Vision Conference 2019 (BMVC \u201919). BMVA Press, 182. Retrieved from https:\/\/bmvc2019.org\/wp-content\/uploads\/papers\/0952-paper.pdf"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2817047"},{"key":"e_1_3_1_42_2","first-page":"413","volume-title":"Proceedings of 16th European Conference on Computer Vision (ECCV \u201920)","author":"Liu Yufan","year":"2020","unstructured":"Yufan Liu, Minglang Qiao, Mai Xu, Bing Li, Weiming Hu, and Ali Borji. 2020. Learning to predict salient faces: A novel visual-audio saliency model. In Proceedings of 16th European Conference on Computer Vision (ECCV \u201920). Springer, 413\u2013429."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3323020"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.04.080"},{"key":"e_1_3_1_45_2","first-page":"7206","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201920)","author":"Martin-Brualla Ricardo","year":"2020","unstructured":"Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi, Jonathan T. Barron, Alexey Dosovitskiy, and Daniel Duckworth. 2020. NeRF in the wild: Neural radiance fields for unconstrained photo collections. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201920), 7206\u20137215."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503250"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00248"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/2996463"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.2966082"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12559-010-9074-z"},{"key":"e_1_3_1_51_2","first-page":"801","volume-title":"Proceedings of International Conference on Computer Vision (ECCV \u201916)","author":"Owens Andrew","year":"2016","unstructured":"Andrew Owens, Jiajun Wu, Josh H. McDermott, William T. Freeman, and Antonio Torralba. 2016. Ambient sound provides supervision for visual learning. In Proceedings of International Conference on Computer Vision (ECCV \u201916). Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling (Eds.), Springer International Publishing, Cham, 801\u2013816."},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","unstructured":"Junting Pan Cristian Canton Kevin McGuinness Noel E. O\u2019Connor Jordi Torres Elisa Sayrol and Xavier and Giro-i Nieto. 2017. SalGAN: Visual saliency prediction with generative adversarial networks. arXiv:1701.01081. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.1701.01081","DOI":"10.48550\/arXiv.1701.01081"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.71"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.3758\/BF03211521"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00886"},{"key":"e_1_3_1_56_2","first-page":"20154","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems","author":"Schwarz Katja","year":"2020","unstructured":"Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger. 2020. GRAF: Generative radiance fields for 3D-aware image synthesis. In Proceedings of the 34th International Conference on Neural Information Processing Systems, 20154\u201320166."},{"key":"e_1_3_1_57_2","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1007\/978-3-319-27674-8_21","volume-title":"Proceedings of International Conference on MultiMedia Modeling","author":"Shen Jialie","year":"2016","unstructured":"Jialie Shen, Liqiang Nie, and Tat-Seng Chua. 2016. Smart ambient sound analysis via structured statistical modeling. In Proceedings of International Conference on MultiMedia Modeling. Qi Tian, Nicu Sebe, Guo-Jun Qi, Benoit Huet, Richang Hong, and Xueliang Liu (Eds.), Springer International Publishing, Cham, 231\u2013243."},{"key":"e_1_3_1_58_2","first-page":"12","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS \u201920)","author":"Sitzmann Vincent","year":"2020","unstructured":"Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell, and Gordon Wetzstein. 2020. Implicit neural representations with periodic activation functions. In Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS \u201920). Curran Associates Inc., Article 626, 12 pages."},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2018.2793599"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01061"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2013.11.028"},{"key":"e_1_3_1_62_2","first-page":"7537","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"33","author":"Tancik Matthew","year":"2020","unstructured":"Matthew Tancik, Pratul Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan Barron, and Ren Ng. 2020. Fourier features let networks learn high-frequency functions in low dimensional domains. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 33, 7537\u20137547."},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","unstructured":"Hamed Rezazadegan Tavakoli Ali Borji Esa Rahtu and Juho Kannala. 2019. DAVE: A deep audio-visual embedding for dynamic saliency prediction. arXiv:abs\/1905.10693. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.1905.10693","DOI":"10.48550\/arXiv.1905.10693"},{"key":"e_1_3_1_64_2","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201918)","author":"Tian Yapeng","year":"2018","unstructured":"Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu. 2018. Audio-visual event localization in unconstrained videos. In Proceedings of the European Conference on Computer Vision (ECCV \u201918)."},{"issue":"1373","key":"e_1_3_1_65_2","doi-asserted-by":"crossref","first-page":"1295","DOI":"10.1098\/rstb.1998.0284","article-title":"Feature binding, attention and object perception","volume":"353","author":"Treisman Anne","year":"1998","unstructured":"Anne Treisman. 1998. Feature binding, attention and object perception. Philosophical Transactions of the Royal Society of London. Series B, Biological Sciences 353, 1373 (1998), 1295\u20131306.","journal-title":"Philosophical Transactions of the Royal Society of London. Series B, Biological Sciences"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1016\/0010-0285(80)90005-5"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00482"},{"issue":"5","key":"e_1_3_1_68_2","doi-asserted-by":"crossref","first-page":"2","DOI":"10.1167\/8.5.2","article-title":"Audiovisual events capture attention: Evidence from temporal order judgments","volume":"8","author":"Burg Erik Van der","year":"2008","unstructured":"Erik Van der Burg, Christian N. L. Olivers, Adelbert W. Bronkhorst, and Jan Theeuwes. 2008. Audiovisual events capture attention: Evidence from temporal order judgments. Journal of Vision 8, 5 (May 2008), 2\u20132.","journal-title":"Journal of Vision"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.358"},{"key":"e_1_3_1_70_2","first-page":"15119","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201921)","author":"Wang Guotao","year":"2021","unstructured":"Guotao Wang, Chenglizhao Chen, Deng-Ping Fan, Aimin Hao, and Hong Qin. 2021. From semantic categories to fixations: A novel weakly-supervised visual-auditory saliency detection approach. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR \u201921), 15119\u201315128."},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2787612"},{"key":"e_1_3_1_72_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00514"},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2924417"},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.1162\/jocn.2008.21173"},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00623"},{"issue":"8","key":"e_1_3_1_76_2","doi-asserted-by":"crossref","first-page":"2163","DOI":"10.1109\/TMM.2019.2947352","article-title":"A dilated inception network for visual saliency prediction","volume":"22","author":"Yang Sheng","year":"2019","unstructured":"Sheng Yang, Guosheng Lin, Qiuping Jiang, and Weisi Lin. 2019. A dilated inception network for visual saliency prediction. IEEE Transactions on Multimedia 22, 8 (2019), 2163\u20132176.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP42928.2021.9506089"},{"key":"e_1_3_1_78_2","volume-title":"Proceedings of the Tenth International Conference on Learning Representations (ICLR)","author":"Yu Sihyun","year":"2022","unstructured":"Sihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim, Junho Kim, Jung-Woo Ha, and Jinwoo Shin. 2022. Generating videos with dynamics-aware implicit generative adversarial networks. In Proceedings of the Tenth International Conference on Learning Representations (ICLR). Retrieved from https:\/\/openreview.net\/forum?id=Czsdv-S4-w9"},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2019.2931735"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","unstructured":"Kai Zhang Gernot Riegler Noah Snavely and Vladlen Koltun. 2020. NeRF++: Analyzing and improving neural radiance fields. arXiv:2010.07492. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2010.07492","DOI":"10.48550\/arXiv.2010.07492"},{"key":"e_1_3_1_81_2","doi-asserted-by":"publisher","DOI":"10.1145\/3576857"},{"key":"e_1_3_1_82_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW54120.2021.00222"},{"key":"e_1_3_1_83_2","doi-asserted-by":"publisher","DOI":"10.5555\/3198485.3198784"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3698881","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3698881","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:09:44Z","timestamp":1750295384000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3698881"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,7]]},"references-count":82,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,4,30]]}},"alternative-id":["10.1145\/3698881"],"URL":"https:\/\/doi.org\/10.1145\/3698881","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,7]]},"assertion":[{"value":"2023-09-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-20","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}