{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T18:26:56Z","timestamp":1776104816758,"version":"3.50.1"},"reference-count":145,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,2,25]],"date-time":"2023-02-25T00:00:00Z","timestamp":1677283200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,8,31]]},"abstract":"<jats:p>\n            Object-level audiovisual saliency detection in 360\u00b0 panoramic real-life dynamic scenes is important for exploring and modeling human perception in immersive environments, also for aiding the development of virtual, augmented, and mixed reality applications in fields such as education, social network, entertainment, and training. To this end, we propose a new task,\n            <jats:bold>p<\/jats:bold>\n            anoramic\n            <jats:bold>a<\/jats:bold>\n            udio\n            <jats:bold>v<\/jats:bold>\n            isual\n            <jats:bold>s<\/jats:bold>\n            alient\n            <jats:bold>o<\/jats:bold>\n            bject\n            <jats:bold>d<\/jats:bold>\n            etection, (\n            <jats:bold>\n              <jats:italic>PAV-SOD<\/jats:italic>\n            <\/jats:bold>\n            <jats:xref ref-type=\"fn\">\n              <jats:sup>1<\/jats:sup>\n            <\/jats:xref>\n            ), which aims to segment the objects grasping most of the human attention in 360\u00b0 panoramic videos reflecting real-life daily scenes. To support the task, we collect\n            <jats:bold>\n              <jats:italic>PAVS10K<\/jats:italic>\n            <\/jats:bold>\n            , the first\n            <jats:bold>p<\/jats:bold>\n            anoramic video dataset for\n            <jats:bold>a<\/jats:bold>\n            udio\n            <jats:bold>v<\/jats:bold>\n            isual\n            <jats:bold>s<\/jats:bold>\n            alient object detection, which consists of 67\u00a04K-resolution equirectangular videos with per-video labels including hierarchical scene categories and associated attributes depicting specific challenges for conducting\n            <jats:italic>PAV-SOD<\/jats:italic>\n            , and 10,465\u00a0uniformly sampled video frames with manually annotated object-level and instance-level pixel-wise masks. The coarse-to-fine annotations enable multi-perspective analysis regarding\n            <jats:italic>PAV-SOD<\/jats:italic>\n            \u00a0modeling. We further systematically benchmark 13\u00a0state-of-the-art\u00a0salient object detection (SOD)\/video object segmentation (VOS) methods based on our\n            <jats:italic>PAVS10K<\/jats:italic>\n            . Besides, we propose a new baseline network, which takes advantage of both visual and audio cues of 360\u00b0 video frames by using a new conditional variational auto-encoder (CVAE). Our\n            <jats:bold>C<\/jats:bold>\n            VAE-based\n            <jats:bold>a<\/jats:bold>\n            udio\n            <jats:bold>v<\/jats:bold>\n            isual\n            <jats:bold>net<\/jats:bold>\n            work, namely,\n            <jats:bold>\n              <jats:italic>CAV-Net<\/jats:italic>\n            <\/jats:bold>\n            , consists of a spatial-temporal visual segmentation network, a convolutional audio-encoding network, and audiovisual distribution estimation modules. As a result, our\n            <jats:italic>CAV-Net<\/jats:italic>\n            \u00a0outperforms all competing models and is able to estimate the aleatoric uncertainties within\n            <jats:italic>PAVS10K<\/jats:italic>\n            . With extensive experimental results, we gain several findings about\n            <jats:italic>PAV-SOD<\/jats:italic>\n            \u00a0challenges and insights towards\n            <jats:italic>PAV-SOD<\/jats:italic>\n            \u00a0model interpretability. We hope that our work could serve as a starting point for advancing SOD towards immersive media.\n          <\/jats:p>","DOI":"10.1145\/3565267","type":"journal-article","created":{"date-parts":[[2022,9,30]],"date-time":"2022-09-30T12:28:25Z","timestamp":1664540905000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["PAV-SOD: A New Task towards Panoramic Audiovisual Saliency Detection"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7110-5948","authenticated-orcid":false,"given":"Yi","family":"Zhang","sequence":"first","affiliation":[{"name":"Univ Rennes, INSA Rennes, CNRS, IETR (UMR 6164), France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1212-1139","authenticated-orcid":false,"given":"Fang-Yi","family":"Chao","sequence":"additional","affiliation":[{"name":"Trinity College Dublin, Ireland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0143-1756","authenticated-orcid":false,"given":"Wassim","family":"Hamidouche","sequence":"additional","affiliation":[{"name":"Univ Rennes, INSA Rennes, CNRS,IETR (UMR 6164), France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0750-0959","authenticated-orcid":false,"given":"Olivier","family":"Deforges","sequence":"additional","affiliation":[{"name":"Univ Rennes, INSA Rennes, CNRS,IETR (UMR 6164), France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,2,25]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3083187.3083218"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2018.2793599"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3083187.3083215"},{"issue":"11","key":"e_1_3_2_5_2","first-page":"2693","article-title":"Predicting head movement in panoramic video: A deep reinforcement learning approach","volume":"41","author":"Xu Mai","year":"2018","unstructured":"Mai Xu, Yuhang Song, Jianyi Wang, MingLang Qiao, Liangyu Huo, and Zulin Wang. 2018. Predicting head movement in panoramic video: A deep reinforcement learning approach. IEEE Trans. Neural Netw. Learn. Syst. 41, 11 (2018), 2693\u20132708.","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"e_1_3_2_6_2","first-page":"488","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Zhang Ziheng","year":"2018","unstructured":"Ziheng Zhang, Yanyu Xu, Jingyi Yu, and Shenghua Gao. 2018. Saliency detection in 360 videos. In Proceedings of the European Conference on Computer Vision (ECCV). 488\u2013503."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00559"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00154"},{"key":"e_1_3_2_9_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Multimedia & Expo Workshops (ICMEW)","author":"Chao Fang-Yi","year":"2020","unstructured":"Fang-Yi Chao, Cagri Ozcinar, Chen Wang, Emin Zerman, Lu Zhang, Wassim Hamidouche, Olivier Deforges, and Aljosa Smolic. 2020. Audio-visual perception of omnidirectional video for virtual reality applications. In Proceedings of the IEEE International Conference on Multimedia & Expo Workshops (ICMEW). IEEE, 1\u20136."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_27"},{"key":"e_1_3_2_11_2","first-page":"638","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Vasudevan Arun Balajee","year":"2020","unstructured":"Arun Balajee Vasudevan, Dengxin Dai, and Luc Van Gool. 2020. Semantic object prediction and spatial sound super-resolution with binaural sounds. In Proceedings of the European Conference on Computer Vision. Springer, 638\u2013655."},{"key":"e_1_3_2_12_2","first-page":"208","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Afouras Triantafyllos","year":"2020","unstructured":"Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman. 2020. Self-supervised learning of audio-visual objects from video. In Proceedings of the European Conference on Computer Vision. Springer, 208\u2013224."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12193-020-00331-1"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2017.2657178"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-019-0149-9"},{"key":"e_1_3_2_16_2","first-page":"12236","article-title":"Few-cost salient object detection with adversarial-paced learning","volume":"33","author":"Zhang Dingwen","year":"2020","unstructured":"Dingwen Zhang, Haibin Tian, and Jungong Han. 2020. Few-cost salient object detection with adversarial-paced learning. Adv. Neural Inf. Process. Syst. 33 (2020), 12236\u201312247.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_17_2","article-title":"Densely nested top-down flows for salient object detection","author":"Fang Chaowei","year":"2021","unstructured":"Chaowei Fang, Haibin Tian, Dingwen Zhang, Qiang Zhang, Jungong Han, and Junwei Han. 2021. Densely nested top-down flows for salient object detection. arXiv preprint arXiv:2102.09133 (2021).","journal-title":"arXiv preprint arXiv:2102.09133"},{"key":"e_1_3_2_18_2","first-page":"186","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Fan Deng-Ping","year":"2018","unstructured":"Deng-Ping Fan, Ming-Ming Cheng, Jiang-Jiang Liu, Shang-Hua Gao, Qibin Hou, and Ali Borji. 2018. Salient objects in clutter: Bringing salient object detection to the foreground. In Proceedings of the European Conference on Computer Vision (ECCV). 186\u2013202."},{"key":"e_1_3_2_19_2","article-title":"Salient object detection in the deep learning era: An in-depth survey","author":"Wang Wenguan","year":"2021","unstructured":"Wenguan Wang, Qiuxia Lai, Huazhu Fu, Jianbing Shen, Haibin Ling, and Ruigang Yang. 2021. Salient object detection in the deep learning era: An in-depth survey. IEEE Trans. Pattern Anal. Mach. Intell. (2021).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_20_2","article-title":"A highly efficient model to study the semantics of salient object detection","author":"Cheng Ming-Ming","year":"2021","unstructured":"Ming-Ming Cheng, Shanghua Gao, Ali Borji, Yong-Qiang Tan, Zheng Lin, and Meng Wang. 2021. A highly efficient model to study the semantics of salient object detection. IEEE Trans. Pattern Anal. Mach. Intell. (2021).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_21_2","first-page":"8554","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Fan Deng-Ping","year":"2019","unstructured":"Deng-Ping Fan, Wenguan Wang, Ming-Ming Cheng, and Jianbing Shen. 2019. Shifting more attention to video salient object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 8554\u20138564."},{"key":"e_1_3_2_22_2","first-page":"16826","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhao Wangbo","year":"2021","unstructured":"Wangbo Zhao, Jing Zhang, Long Li, Nick Barnes, Nian Liu, and Junwei Han. 2021. Weakly supervised video salient object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 16826\u201316835."},{"key":"e_1_3_2_23_2","first-page":"1553","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhang Miao","year":"2021","unstructured":"Miao Zhang, Jie Liu, Yifei Wang, Yongri Piao, Shunyu Yao, Wei Ji, Jingjing Li, Huchuan Lu, and Zhongxuan Luo. 2021. Dynamic context-sensitive filtering network for video salient object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 1553\u20131563."},{"key":"e_1_3_2_24_2","article-title":"Global-and-local collaborative learning for co-salient object detection","author":"Cong Runmin","year":"2022","unstructured":"Runmin Cong, Ning Yang, Chongyi Li, Huazhu Fu, Yao Zhao, Qingming Huang, and Sam Kwong. 2022. Global-and-local collaborative learning for co-salient object detection. arXiv preprint arXiv:2204.08917 (2022).","journal-title":"arXiv preprint arXiv:2204.08917"},{"key":"e_1_3_2_25_2","first-page":"12288","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Fan Qi","year":"2021","unstructured":"Qi Fan, Deng-Ping Fan, Huazhu Fu, Chi-Keung Tang, Ling Shao, and Yu-Wing Tai. 2021. Group collaborative learning for co-salient object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 12288\u201312298."},{"key":"e_1_3_2_26_2","first-page":"8846","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Hsu Kuang-Jui","year":"2019","unstructured":"Kuang-Jui Hsu, Yen-Yu Lin, and Yung-Yu Chuang. 2019. DeepCO3: Deep instance co-segmentation by co-peak search and co-saliency detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 8846\u20138855."},{"key":"e_1_3_2_27_2","first-page":"6959","article-title":"CoADNet: Collaborative aggregation-and-distribution networks for co-salient object detection","volume":"33","author":"Zhang Qijian","year":"2020","unstructured":"Qijian Zhang, Runmin Cong, Junhui Hou, Chongyi Li, and Yao Zhao. 2020. CoADNet: Collaborative aggregation-and-distribution networks for co-salient object detection. Adv. Neural Inf. Process. Syst. 33 (2020), 6959\u20136970.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3028289"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58610-2_17"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i2.16191"},{"key":"e_1_3_2_31_2","article-title":"MobileSal: Extremely efficient RGB-D salient object detection","author":"Wu Yu-Huan","year":"2021","unstructured":"Yu-Huan Wu, Yun Liu, Jun Xu, Jia-Wang Bian, Yu-Chao Gu, and Ming-Ming Cheng. 2021. MobileSal: Extremely efficient RGB-D salient object detection. IEEE Trans. Pattern Anal. Mach. Intell. (2021).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"1","key":"e_1_3_2_32_2","first-page":"160","article-title":"RGB-T image saliency detection via collaborative graph learning","volume":"22","author":"Tu Zhengzheng","year":"2019","unstructured":"Zhengzheng Tu, Tian Xia, Chenglong Li, Xiaoxiao Wang, Yan Ma, and Jin Tang. 2019. RGB-T image saliency detection via collaborative graph learning. IEEE Trans. Multim. 22, 1 (2019), 160\u2013173.","journal-title":"IEEE Trans. Multim."},{"issue":"12","key":"e_1_3_2_33_2","doi-asserted-by":"crossref","first-page":"4421","DOI":"10.1109\/TCSVT.2019.2951621","article-title":"RGBT salient object detection: Benchmark and a novel cooperative ranking approach","volume":"30","author":"Tang Jin","year":"2019","unstructured":"Jin Tang, Dongzhe Fan, Xiaoxiao Wang, Zhengzheng Tu, and Chenglong Li. 2019. RGBT salient object detection: Benchmark and a novel cooperative ranking approach. IEEE Trans. Circ. Syst. Vid. Technol. 30, 12 (2019), 4421\u20134433.","journal-title":"IEEE Trans. Circ. Syst. Vid. Technol."},{"issue":"5","key":"e_1_3_2_34_2","doi-asserted-by":"crossref","first-page":"1804","DOI":"10.1109\/TCSVT.2020.3014663","article-title":"Revisiting feature fusion for RGB-T salient object detection","volume":"31","author":"Zhang Qiang","year":"2020","unstructured":"Qiang Zhang, Tonglin Xiao, Nianchang Huang, Dingwen Zhang, and Jungong Han. 2020. Revisiting feature fusion for RGB-T salient object detection. IEEE Trans. Circ. Syst. Vid. Technol. 31, 5 (2020), 1804\u20131818.","journal-title":"IEEE Trans. Circ. Syst. Vid. Technol."},{"key":"e_1_3_2_35_2","first-page":"2806","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Li Nianyi","year":"2014","unstructured":"Nianyi Li, Jinwei Ye, Yu Ji, Haibin Ling, and Jingyi Yu. 2014. Saliency detection on light field. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2806\u20132813."},{"key":"e_1_3_2_36_2","article-title":"Memory-oriented decoder for light field salient object detection","volume":"32","author":"Zhang Miao","year":"2019","unstructured":"Miao Zhang, Jingjing Li, Ji Wei, Yongri Piao, and Huchuan Lu. 2019. Memory-oriented decoder for light field salient object detection. Adv. Neural Inf. Process. Syst. 32 (2019).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_37_2","article-title":"Learning synergistic attention for light field salient object detection","author":"Zhang Yi","year":"2021","unstructured":"Yi Zhang, Geng Chen, Qian Chen, Yujia Sun, Yong Xia, Olivier Deforges, Wassim Hamidouche, and Lu Zhang. 2021. Learning synergistic attention for light field salient object detection. arXiv preprint arXiv:2104.13916 (2021).","journal-title":"arXiv preprint arXiv:2104.13916"},{"key":"e_1_3_2_38_2","first-page":"7234","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zeng Yi","year":"2019","unstructured":"Yi Zeng, Pingping Zhang, Jianming Zhang, Zhe Lin, and Huchuan Lu. 2019. Towards high-resolution salient object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 7234\u20137243."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3045624"},{"key":"e_1_3_2_40_2","article-title":"RRNet: Relational reasoning network with parallel multi-scale attention for salient object detection in optical remote sensing images","author":"Cong Runmin","year":"2021","unstructured":"Runmin Cong, Yumo Zhang, Leyuan Fang, Jun Li, Chunjie Zhang, Yao Zhao, and Sam Kwong. 2021. RRNet: Relational reasoning network with parallel multi-scale attention for salient object detection in optical remote sensing images. IEEE Trans. Geosci. Rem. Sens. (2021).","journal-title":"IEEE Trans. Geosci. Rem. Sens."},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2019.2925070"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3042084"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.6867718"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1152\/jn.1986.56.3.640"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00482"},{"key":"e_1_3_2_46_2","first-page":"15119","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Guotao","year":"2021","unstructured":"Guotao Wang, Chenglizhao Chen, Deng-Ping Fan, Aimin Hao, and Hong Qin. 2021. From semantic categories to fixations: A novel weakly-supervised visual-auditory saliency detection approach. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 15119\u201315128."},{"key":"e_1_3_2_47_2","first-page":"510","volume-title":"Proceedings of the International Conference on Intelligent Computing","author":"Cheng Shuaiyang","year":"2021","unstructured":"Shuaiyang Cheng, Liang Song, Jingjing Tang, and Shihui Guo. 2021. Audio-visual salient object detection. In Proceedings of the International Conference on Intelligent Computing. Springer, 510\u2013521."},{"key":"e_1_3_2_48_2","first-page":"3458","volume-title":"Proceedings of the IEEE International Conference on Image Processing (ICIP)","author":"Zhang Yi","year":"2020","unstructured":"Yi Zhang, Lu Zhang, Wassim Hamidouche, and Olivier Deforges. 2020. A fixation-based 360 benchmark dataset for salient object detection. In Proceedings of the IEEE International Conference on Image Processing (ICIP). 3458\u20133462."},{"issue":"1","key":"e_1_3_2_49_2","first-page":"38","article-title":"Distortion-adaptive salient object detection in 360\u00b0 omnidirectional images","volume":"14","author":"Li Jia","year":"2019","unstructured":"Jia Li, Jinming Su, Changqun Xia, and Yonghong Tian. 2019. Distortion-adaptive salient object detection in 360\u00b0 omnidirectional images. IEEE J. Select. Topics Sig. Process. 14, 1 (2019), 38\u201348.","journal-title":"IEEE J. Select. Topics Sig. Process."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2020.3023636"},{"key":"e_1_3_2_51_2","article-title":"Learning structured output representation using deep conditional generative models","volume":"28","author":"Sohn Kihyuk","year":"2015","unstructured":"Kihyuk Sohn, Honglak Lee, and Xinchen Yan. 2015. Learning structured output representation using deep conditional generative models. Adv. Neural Inf. Process. Syst. 28 (2015).","journal-title":"Adv. Neural Inf. Process. Syst."},{"issue":"1","key":"e_1_3_2_52_2","first-page":"220","article-title":"Revisiting video saliency prediction in the deep learning era","volume":"43","author":"Wang Wenguan","year":"2019","unstructured":"Wenguan Wang, Jianbing Shen, Jianwen Xie, Ming-Ming Cheng, Haibin Ling, and Ali Borji. 2019. Revisiting video saliency prediction in the deep learning era. IEEE Trans. Neural Netw. Learn. Syst. 43, 1 (2019), 220\u2013237.","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_35"},{"key":"e_1_3_2_54_2","first-page":"776","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Gemmeke Jort F.","year":"2017","unstructured":"Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Audio set: An ontology and human-labeled dataset for audio events. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 776\u2013780."},{"key":"e_1_3_2_55_2","first-page":"247","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Tian Yapeng","year":"2018","unstructured":"Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu. 2018. Audio-visual event localization in unconstrained videos. In Proceedings of the European Conference on Computer Vision (ECCV). 247\u2013263."},{"key":"e_1_3_2_56_2","first-page":"721","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Chen Honglie","year":"2020","unstructured":"Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman. 2020. VGGSound: A large-scale audio-visual dataset. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 721\u2013725."},{"key":"e_1_3_2_57_2","article-title":"ObjectFolder: A dataset of objects with implicit visual, auditory, and tactile representations","author":"Gao Ruohan","year":"2021","unstructured":"Ruohan Gao, Yen-Yu Chang, Shivani Mall, Li Fei-Fei, and Jiajun Wu. 2021. ObjectFolder: A dataset of objects with implicit visual, auditory, and tactile representations. arXiv preprint arXiv:2109.07991 (2021).","journal-title":"arXiv preprint arXiv:2109.07991"},{"key":"e_1_3_2_58_2","article-title":"Audio-visual synchronisation in the wild","author":"Chen Honglie","year":"2021","unstructured":"Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman. 2021. Audio-visual synchronisation in the wild. arXiv preprint arXiv:2112.04432 (2021).","journal-title":"arXiv preprint arXiv:2112.04432"},{"key":"e_1_3_2_59_2","first-page":"9701","volume-title":"Proceedings of the IEEE International Conference on Robotics and Automation (ICRA)","author":"Gan Chuang","year":"2020","unstructured":"Chuang Gan, Yiwei Zhang, Jiajun Wu, Boqing Gong, and Joshua B. Tenenbaum. 2020. Look, listen, and act: Towards audio-visual embodied navigation. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). IEEE, 9701\u20139707."},{"key":"e_1_3_2_60_2","first-page":"17","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Chen Changan","year":"2020","unstructured":"Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc Amengual Gari, Ziad Al-Halah, Vamsi Krishna Ithapu, Philip Robinson, and Kristen Grauman. 2020. SoundSpaces: Audio-visual navigation in 3D environments. In Proceedings of the European Conference on Computer Vision. Springer, 17\u201336."},{"key":"e_1_3_2_61_2","article-title":"Self-supervised generation of spatial audio for 360 video","volume":"31","author":"Morgado Pedro","year":"2018","unstructured":"Pedro Morgado, Nuno Nvasconcelos, Timothy Langlois, and Oliver Wang. 2018. Self-supervised generation of spatial audio for 360 video. Adv. Neural Inf. Process. Syst. 31 (2018).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_62_2","article-title":"Geometry-aware multi-task learning for binaural audio generation from video","author":"Garg Rishabh","year":"2021","unstructured":"Rishabh Garg, Ruohan Gao, and Kristen Grauman. 2021. Geometry-aware multi-task learning for binaural audio generation from video. arXiv preprint arXiv:2111.10882 (2021).","journal-title":"arXiv preprint arXiv:2111.10882"},{"key":"e_1_3_2_63_2","first-page":"3761","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Narasimhan Medhini","year":"2022","unstructured":"Medhini Narasimhan, Shiry Ginosar, Andrew Owens, Alexei A. Efros, and Trevor Darrell. 2022. Strumming to the beat: Audio-conditioned contrastive video textures. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 3761\u20133770."},{"key":"e_1_3_2_64_2","first-page":"15516","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Changan","year":"2021","unstructured":"Changan Chen, Ziad Al-Halah, and Kristen Grauman. 2021. Semantic audio-visual navigation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 15516\u201315525."},{"key":"e_1_3_2_65_2","first-page":"275","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Majumder Sagnik","year":"2021","unstructured":"Sagnik Majumder, Ziad Al-Halah, and Kristen Grauman. 2021. Move2Hear: Active audio-visual source separation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 275\u2013285."},{"key":"e_1_3_2_66_2","first-page":"2357","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Rouditchenko Andrew","year":"2019","unstructured":"Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh McDermott, and Antonio Torralba. 2019. Self-supervised audio-visual co-segmentation. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2357\u20132361."},{"key":"e_1_3_2_67_2","first-page":"16867","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Honglie","year":"2021","unstructured":"Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman. 2021. Localizing visual sounds the hard way. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 16867\u201316876."},{"key":"e_1_3_2_68_2","first-page":"3879","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Gao Ruohan","year":"2019","unstructured":"Ruohan Gao and Kristen Grauman. 2019. Co-separating sounds of visual objects. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 3879\u20133888."},{"key":"e_1_3_2_69_2","article-title":"Visual sound localization in the wild by cross-modal interference erasing","author":"Liu Xian","year":"2022","unstructured":"Xian Liu, Rui Qian, Hang Zhou, Di Hu, Weiyao Lin, Ziwei Liu, Bolei Zhou, and Xiaowei Zhou. 2022. Visual sound localization in the wild by cross-modal interference erasing. arXiv preprint arXiv:2202.06406 (2022).","journal-title":"arXiv preprint arXiv:2202.06406"},{"key":"e_1_3_2_70_2","first-page":"2745","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Tian Yapeng","year":"2021","unstructured":"Yapeng Tian, Di Hu, and Chenliang Xu. 2021. Cyclic co-learning of sounding object visual grounding and sound separation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2745\u20132754."},{"key":"e_1_3_2_71_2","first-page":"10077","article-title":"Discriminative sounding objects localization via self-supervised audiovisual matching","volume":"33","author":"Hu Di","year":"2020","unstructured":"Di Hu, Rui Qian, Minyue Jiang, Xiao Tan, Shilei Wen, Errui Ding, Weiyao Lin, and Dejing Dou. 2020. Discriminative sounding objects localization via self-supervised audiovisual matching. Adv. Neural Inf. Process. Syst. 33 (2020), 10077\u201310087.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_72_2","article-title":"Class-aware sounding objects localization via audiovisual correspondence","author":"Hu Di","year":"2021","unstructured":"Di Hu, Yake Wei, Rui Qian, Weiyao Lin, Ruihua Song, and Ji-Rong Wen. 2021. Class-aware sounding objects localization via audiovisual correspondence. IEEE Trans. Pattern Anal. Mach. Intell. (2021).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_73_2","article-title":"Self-supervised object detection from audio-visual correspondence","author":"Afouras Triantafyllos","year":"2021","unstructured":"Triantafyllos Afouras, Yuki M. Asano, Francois Fagan, Andrea Vedaldi, and Florian Metze. 2021. Self-supervised object detection from audio-visual correspondence. arXiv preprint arXiv:2104.06401 (2021).","journal-title":"arXiv preprint arXiv:2104.06401"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1145\/3204949.3208139"},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/3304109.3325820"},{"key":"e_1_3_2_76_2","first-page":"932","volume-title":"Proceedings of the 26th ACM International Conference on Multimedia","author":"Li Chen","year":"2018","unstructured":"Chen Li, Mai Xu, Xinzhe Du, and Zulin Wang. 2018. Bridge the gap between VQA and human behavior on omnidirectional video: A large-scale dataset and a deep learning model. In Proceedings of the 26th ACM International Conference on Multimedia. 932\u2013940."},{"key":"e_1_3_2_77_2","doi-asserted-by":"crossref","first-page":"1007","DOI":"10.1145\/3343031.3350947","volume-title":"Proceedings of the 27th ACM International Conference on Multimedia","author":"Agtzidis Ioannis","year":"2019","unstructured":"Ioannis Agtzidis, Mikhail Startsev, and Michael Dorr. 2019. 360-degree video gaze behaviour: A ground-truth data set and a classification algorithm for eye movements. In Proceedings of the 27th ACM International Conference on Multimedia. 1007\u20131015."},{"key":"e_1_3_2_78_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Multimedia & Expo Workshops (ICMEW)","author":"Chao Fang-Yi","year":"2018","unstructured":"Fang-Yi Chao, Lu Zhang, Wassim Hamidouche, and Olivier Deforges. 2018. SalGAN360: Visual saliency prediction on 360 degree images with generative adversarial networks. In Proceedings of the IEEE International Conference on Multimedia & Expo Workshops (ICMEW). IEEE, 1\u20134."},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.2987682"},{"key":"e_1_3_2_80_2","first-page":"305","volume-title":"Proceedings of the International Conference on Pattern Recognition","author":"Dahou Yasser","year":"2021","unstructured":"Yasser Dahou, Marouane Tliba, Kevin McGuinness, and Noel O\u2019Connor. 2021. ATSal: An attention based architecture for saliency prediction in 360\u00b0 videos. In Proceedings of the International Conference on Pattern Recognition. Springer, 305\u2013320."},{"issue":"9","key":"e_1_3_2_81_2","first-page":"2331","article-title":"The prediction of saliency map for head and eye movements in 360 degree images","volume":"22","author":"Zhu Yucheng","year":"2019","unstructured":"Yucheng Zhu, Guangtao Zhai, Xiongkuo Min, and Jiantao Zhou. 2019. The prediction of saliency map for head and eye movements in 360 degree images. IEEE Trans. Multim. 22, 9 (2019), 2331\u20132344.","journal-title":"IEEE Trans. Multim."},{"key":"e_1_3_2_82_2","first-page":"3750","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Djilali Yasser Abdelaziz Dahou","year":"2021","unstructured":"Yasser Abdelaziz Dahou Djilali, Kevin McGuinness, and Noel E. O\u2019Connor. 2021. Simple baselines can fool 360deg saliency metrics. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 3750\u20133756."},{"key":"e_1_3_2_83_2","article-title":"DAVE: A deep audio-visual embedding for dynamic saliency prediction","author":"Tavakoli Hamed R.","year":"2019","unstructured":"Hamed R. Tavakoli, Ali Borji, Esa Rahtu, and Juho Kannala. 2019. DAVE: A deep audio-visual embedding for dynamic saliency prediction. arXiv preprint arXiv:1905.10693 (2019).","journal-title":"arXiv preprint arXiv:1905.10693"},{"key":"e_1_3_2_84_2","article-title":"SoundNet: Learning sound representations from unlabeled video","volume":"29","author":"Aytar Yusuf","year":"2016","unstructured":"Yusuf Aytar, Carl Vondrick, and Antonio Torralba. 2016. SoundNet: Learning sound representations from unlabeled video. Adv. Neural Inf. Process. Syst. 29 (2016).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_85_2","first-page":"413","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Liu Yufan","year":"2020","unstructured":"Yufan Liu, Minglang Qiao, Mai Xu, Bing Li, Weiming Hu, and Ali Borji. 2020. Learning to predict salient faces: A novel visual-audio saliency model. In Proceedings of the European Conference on Computer Vision. Springer, 413\u2013429."},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.2966082"},{"key":"e_1_3_2_87_2","first-page":"1604","volume-title":"Proceedings of the IEEE International Conference on Image Processing (ICIP)","author":"Yao Shunyu","year":"2021","unstructured":"Shunyu Yao, Xiongkuo Min, and Guangtao Zhai. 2021. Deep audio-visual fusion neural network for saliency estimation. In Proceedings of the IEEE International Conference on Image Processing (ICIP). 1604\u20131608."},{"key":"e_1_3_2_88_2","first-page":"3520","volume-title":"Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"Jain Samyak","year":"2020","unstructured":"Samyak Jain, Pradeep Yarlagadda, Shreyank Jyoti, Shyamgopal Karthik, Ramanathan Subramanian, and Vineet Gandhi. 2020. ViNet: Pushing the limits of visual modality for audio-visual saliency prediction. In Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS). 3520\u20133527."},{"key":"e_1_3_2_89_2","article-title":"GASP: Gated attention for saliency prediction","author":"Abawi Fares","year":"2021","unstructured":"Fares Abawi, Tom Weber, and Stefan Wermter. 2021. GASP: Gated attention for saliency prediction. In Proceedings of the 30th International Joint Conference on Artificial Intelligence.","journal-title":"Proceedings of the 30th International Joint Conference on Artificial Intelligence"},{"issue":"4","key":"e_1_3_2_90_2","first-page":"487","article-title":"A biologically motivated, proto-object-based audiovisual saliency model","volume":"1","author":"Ramenahalli Sudarshan","year":"2020","unstructured":"Sudarshan Ramenahalli. 2020. A biologically motivated, proto-object-based audiovisual saliency model. Artif. Intell. 1, 4 (2020), 487\u2013509.","journal-title":"Artif. Intell."},{"key":"e_1_3_2_91_2","first-page":"355","volume-title":"Proceedings of the IEEE International Conference on Visual Communications and Image Processing (VCIP)","author":"Chao Fang-Yi","year":"2020","unstructured":"Fang-Yi Chao, Cagri Ozcinar, Lu Zhang, Wassim Hamidouche, Olivier Deforges, and Aljosa Smolic. 2020. Towards audio-visual saliency prediction for omnidirectional video with spatial audio. In Proceedings of the IEEE International Conference on Visual Communications and Image Processing (VCIP). IEEE, 355\u2013358."},{"key":"e_1_3_2_92_2","first-page":"1","volume-title":"Proceedings of the 17th International Conference on Machine Vision and Applications (MVA)","author":"Cokelek Mert","year":"2021","unstructured":"Mert Cokelek, Nevrez Imamoglu, Cagri Ozcinar, Erkut Erdem, and Aykut Erdem. 2021. Leveraging frequency based salient spatial sound localization to improve 360\u00b0 video saliency prediction. In Proceedings of the 17th International Conference on Machine Vision and Applications (MVA). 1\u20135."},{"key":"e_1_3_2_93_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.85"},{"key":"e_1_3_2_94_2","first-page":"2192","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Li Fuxin","year":"2013","unstructured":"Fuxin Li, Taeyoung Kim, Ahmad Humayun, David Tsai, and James M. Rehg. 2013. Video segmentation by tracking many figure-ground segments. In Proceedings of the IEEE International Conference on Computer Vision. 2192\u20132199."},{"issue":"6","key":"e_1_3_2_95_2","doi-asserted-by":"crossref","first-page":"1187","DOI":"10.1109\/TPAMI.2013.242","article-title":"Segmentation of moving objects by long term video analysis","volume":"36","author":"Ochs Peter","year":"2013","unstructured":"Peter Ochs, Jitendra Malik, and Thomas Brox. 2013. Segmentation of moving objects by long term video analysis. IEEE Trans. Pattern Anal. Mach. Intell. 36, 6 (2013), 1187\u20131200.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_96_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2015.2425544"},{"key":"e_1_3_2_97_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2015.2460013"},{"issue":"12","key":"e_1_3_2_98_2","first-page":"2527","article-title":"Saliency detection for unconstrained videos using superpixel-level graph and spatiotemporal propagation","volume":"27","author":"Liu Zhi","year":"2016","unstructured":"Zhi Liu, Junhao Li, Linwei Ye, Guangling Sun, and Liquan Shen. 2016. Saliency detection for unconstrained videos using superpixel-level graph and spatiotemporal propagation. IEEE Trans. Circ. Syst. Vid. Technol. 27, 12 (2016), 2527\u20132542.","journal-title":"IEEE Trans. Circ. Syst. Vid. Technol."},{"issue":"1","key":"e_1_3_2_99_2","first-page":"349","article-title":"A benchmark dataset and saliency-guided stacked autoencoders for video-based salient object detection","volume":"27","author":"Li Jia","year":"2017","unstructured":"Jia Li, Changqun Xia, and Xiaowu Chen. 2017. A benchmark dataset and saliency-guided stacked autoencoders for video-based salient object detection. IEEE Trans. Image Process. 27, 1 (2017), 349\u2013364.","journal-title":"IEEE Trans. Image Process."},{"key":"e_1_3_2_100_2","first-page":"7284","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Yan Pengxiang","year":"2019","unstructured":"Pengxiang Yan, Guanbin Li, Yuan Xie, Zhen Li, Chuan Wang, Tianshui Chen, and Liang Lin. 2019. Semi-supervised video salient object detection using pseudo-labels. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 7284\u20137293."},{"key":"e_1_3_2_101_2","first-page":"10869","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Gu Yuchao","year":"2020","unstructured":"Yuchao Gu, Lijuan Wang, Ziqin Wang, Yun Liu, Ming-Ming Cheng, and Shao-Ping Lu. 2020. Pyramid constrained self-attention network for fast video salient object detection. In Proceedings of the AAAI Conference on Artificial Intelligence. 10869\u201310876."},{"key":"e_1_3_2_102_2","first-page":"1200","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Schmidt Christian","year":"2022","unstructured":"Christian Schmidt, Ali Athar, Sabarinath Mahadevan, and Bastian Leibe. 2022. D2Conv3D: Dynamic dilated convolutions for object segmentation in videos. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 1200\u20131209."},{"key":"e_1_3_2_103_2","first-page":"15455","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ren Sucheng","year":"2021","unstructured":"Sucheng Ren, Wenxi Liu, Yongtuo Liu, Haoxin Chen, Guoqiang Han, and Shengfeng He. 2021. Reciprocal transformations for unsupervised video object segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 15455\u201315464."},{"key":"e_1_3_2_104_2","article-title":"Making a case for 3D convolutions for object segmentation in videos","author":"Mahadevan Sabarinath","year":"2020","unstructured":"Sabarinath Mahadevan, Ali Athar, Aljo\u0161a O\u0161ep, Sebastian Hennen, Laura Leal-Taix\u00e9, and Bastian Leibe. 2020. Making a case for 3D convolutions for object segmentation in videos. arXiv preprint arXiv:2008.11516 (2020).","journal-title":"arXiv preprint arXiv:2008.11516"},{"key":"e_1_3_2_105_2","first-page":"4922","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Ji Ge-Peng","year":"2021","unstructured":"Ge-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan, Jianbing Shen, and Ling Shao. 2021. Full-duplex strategy for video object segmentation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4922\u20134933."},{"key":"e_1_3_2_106_2","first-page":"8781","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhang Kaihua","year":"2021","unstructured":"Kaihua Zhang, Zicheng Zhao, Dong Liu, Qingshan Liu, and Bo Liu. 2021. Deep transport network for unsupervised video object segmentation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 8781\u20138790."},{"key":"e_1_3_2_107_2","first-page":"13066","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Zhou Tianfei","year":"2020","unstructured":"Tianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao, Jianwu Li, and Ling Shao. 2020. Motion-attentive transition for zero-shot video object segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence. 13066\u201313073."},{"key":"e_1_3_2_108_2","first-page":"3623","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lu Xiankai","year":"2019","unstructured":"Xiankai Lu, Wenguan Wang, Chao Ma, Jianbing Shen, Ling Shao, and Fatih Porikli. 2019. See more, know more: Unsupervised video object segmentation with co-attention siamese networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3623\u20133632."},{"key":"e_1_3_2_109_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2020.3028192"},{"key":"e_1_3_2_110_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_3_2_111_2","first-page":"1396","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Hu Hou-Ning","year":"2017","unstructured":"Hou-Ning Hu, Yen-Chen Lin, Ming-Yu Liu, Hsien-Tzu Cheng, Yung-Ju Chang, and Min Sun. 2017. Deep 360 pilot: Learning a deep agent for piloting through 360 sports videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 1396\u20131405."},{"key":"e_1_3_2_112_2","first-page":"12959","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Zhao Pengyu","year":"2020","unstructured":"Pengyu Zhao, Ansheng You, Yuanxing Zhang, Jiaying Liu, Kaigui Bian, and Yunhai Tong. 2020. Spherical criteria for fast and accurate 360 object detection. In Proceedings of the AAAI Conference on Artificial Intelligence. 12959\u201312966."},{"key":"e_1_3_2_113_2","first-page":"3642","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Wang Kuan-Hsun","year":"2019","unstructured":"Kuan-Hsun Wang and Shang-Hong Lai. 2019. Object detection in curved space for 360-degree camera. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 3642\u20133646."},{"key":"e_1_3_2_114_2","first-page":"18","volume-title":"Proceedings of the International Conference on Virtual Reality and Visualization (ICVRV)","author":"Zhang Yiming","year":"2017","unstructured":"Yiming Zhang, Xiangyun Xiao, and Xubo Yang. 2017. Real-time object detection for 360-degree panoramic image using CNN. In Proceedings of the International Conference on Virtual Reality and Visualization (ICVRV). IEEE, 18\u201323."},{"key":"e_1_3_2_115_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.153"},{"key":"e_1_3_2_116_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.407"},{"key":"e_1_3_2_117_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.43"},{"key":"e_1_3_2_118_2","first-page":"5455","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Li Guanbin","year":"2015","unstructured":"Guanbin Li and Yizhou Yu. 2015. Visual saliency based on multiscale deep features. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5455\u20135463."},{"key":"e_1_3_2_119_2","first-page":"136","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Lijun","year":"2017","unstructured":"Lijun Wang, Huchuan Lu, Yifan Wang, Mengyang Feng, Dong Wang, Baocai Yin, and Xiang Ruan. 2017. Learning to detect salient objects with image-level supervision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 136\u2013145."},{"key":"e_1_3_2_120_2","first-page":"2386","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Li Guanbin","year":"2017","unstructured":"Guanbin Li, Yuan Xie, Liang Lin, and Yizhou Yu. 2017. Instance-level salient object segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2386\u20132395."},{"key":"e_1_3_2_121_2","first-page":"14769","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhang Kaihao","year":"2021","unstructured":"Kaihao Zhang, Dongxu Li, Wenhan Luo, Wenqi Ren, Bj\u00f6rn Stenger, Wei Liu, Hongdong Li, and Ming-Hsuan Yang. 2021. Benchmarking ultra-high-definition image super-resolution. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 14769\u201314778."},{"key":"e_1_3_2_122_2","first-page":"14030","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Deng Senyou","year":"2021","unstructured":"Senyou Deng, Wenqi Ren, Yanyang Yan, Tao Wang, Fenglong Song, and Xiaochun Cao. 2021. Multi-scale separable network for ultra-high-definition video deblurring. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 14030\u201314039."},{"key":"e_1_3_2_123_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2020.2966864"},{"key":"e_1_3_2_124_2","doi-asserted-by":"publisher","DOI":"10.1145\/3329119"},{"key":"e_1_3_2_125_2","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang Jing","year":"2020","unstructured":"Jing Zhang, Deng-Ping Fan, Yuchao Dai, Saeed Anwar, Fatemeh Sadat Saleh, Tong Zhang, and Nick Barnes. 2020. UC-Net: Uncertainty inspired RGB-D saliency detection via conditional variational autoencoders. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_126_2","article-title":"Dense uncertainty estimation","author":"Zhang Jing","year":"2021","unstructured":"Jing Zhang, Yuchao Dai, Mochu Xiang, Deng-Ping Fan, Peyman Moghadam, Mingyi He, Christian Walder, Kaihao Zhang, Mehrtash Harandi, and Nick Barnes. 2021. Dense uncertainty estimation. arXiv preprint arXiv:2110.06427 (2021).","journal-title":"arXiv preprint arXiv:2110.06427"},{"key":"e_1_3_2_127_2","article-title":"Auto-encoding variational Bayes","author":"Kingma Diederik P.","year":"2013","unstructured":"Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational Bayes. arXiv preprint arXiv:1312.6114 (2013).","journal-title":"arXiv preprint arXiv:1312.6114"},{"key":"e_1_3_2_128_2","first-page":"12179","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Ranftl Ren\u00e9","year":"2021","unstructured":"Ren\u00e9 Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. 2021. Vision transformers for dense prediction. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). 12179\u201312188."},{"key":"e_1_3_2_129_2","article-title":"An image is worth 16x16 words: Transformers for image recognition at scale","author":"Dosovitskiy Alexey","year":"2020","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et\u00a0al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020).","journal-title":"arXiv preprint arXiv:2010.11929"},{"key":"e_1_3_2_130_2","first-page":"12321","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Wei Jun","year":"2020","unstructured":"Jun Wei, Shuhui Wang, and Qingming Huang. 2020. F \\(^3\\) Net: Fusion, feedback and focus for salient object detection. In Proceedings of the AAAI Conference on Artificial Intelligence. 12321\u201312328."},{"key":"e_1_3_2_131_2","article-title":"Adam: A method for stochastic optimization","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014).","journal-title":"arXiv preprint arXiv:1412.6980"},{"key":"e_1_3_2_132_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00403"},{"key":"e_1_3_2_133_2","first-page":"7264","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Wu Zhe","year":"2019","unstructured":"Zhe Wu, Li Su, and Qingming Huang. 2019. Stacked cross refinement network for edge-aware salient object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 7264\u20137273."},{"key":"e_1_3_2_134_2","first-page":"9413","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Pang Youwei","year":"2020","unstructured":"Youwei Pang, Xiaoqi Zhao, Lihe Zhang, and Huchuan Lu. 2020. Multi-scale interactive network for salient object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 9413\u20139422."},{"key":"e_1_3_2_135_2","first-page":"13025","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wei Jun","year":"2020","unstructured":"Jun Wei, Shuhui Wang, Zhe Wu, Chi Su, Qingming Huang, and Qi Tian. 2020. Label decoupling framework for salient object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 13025\u201313034."},{"key":"e_1_3_2_136_2","first-page":"702","volume-title":"European Conference on Computer Vision","author":"Gao Shang-Hua","year":"2020","unstructured":"Shang-Hua Gao, Yong-Qiang Tan, Ming-Ming Cheng, Chengze Lu, Yunpeng Chen, and Shuicheng Yan. 2020. Highly efficient salient object detection with 100k parameters. In European Conference on Computer Vision. 702\u2013721."},{"key":"e_1_3_2_137_2","first-page":"35","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Zhao Xiaoqi","year":"2020","unstructured":"Xiaoqi Zhao, Youwei Pang, Lihe Zhang, Huchuan Lu, and Lei Zhang. 2020. Suppress and balance: A simple gated network for salient object detection. In Proceedings of the European Conference on Computer Vision. Springer, 35\u201351."},{"key":"e_1_3_2_138_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206596"},{"key":"e_1_3_2_139_2","first-page":"733","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Perazzi Federico","year":"2012","unstructured":"Federico Perazzi, Philipp Kr\u00e4henb\u00fchl, Yael Pritch, and Alexander Hornung. 2012. Saliency filters: Contrast based filtering for salient region detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 733\u2013740."},{"key":"e_1_3_2_140_2","first-page":"4548","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Fan Deng-Ping","year":"2017","unstructured":"Deng-Ping Fan, Ming-Ming Cheng, Yun Liu, Tao Li, and Ali Borji. 2017. Structure-measure: A new way to evaluate foreground maps. In Proceedings of the IEEE International Conference on Computer Vision. 4548\u20134557."},{"key":"e_1_3_2_141_2","article-title":"Enhanced-alignment measure for binary foreground map evaluation","author":"Fan Deng-Ping","year":"2018","unstructured":"Deng-Ping Fan, Cheng Gong, Yang Cao, Bo Ren, Ming-Ming Cheng, and Ali Borji. 2018. Enhanced-alignment measure for binary foreground map evaluation. arXiv preprint arXiv:1805.10421 (2018).","journal-title":"arXiv preprint arXiv:1805.10421"},{"key":"e_1_3_2_142_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"issue":"2","key":"e_1_3_2_143_2","first-page":"652","article-title":"Res2Net: A new multi-scale backbone architecture","volume":"43","author":"Gao Shang-Hua","year":"2019","unstructured":"Shang-Hua Gao, Ming-Ming Cheng, Kai Zhao, Xin-Yu Zhang, Ming-Hsuan Yang, and Philip Torr. 2019. Res2Net: A new multi-scale backbone architecture. IEEE Trans. Pattern Anal. Mach. Intell. 43, 2 (2019), 652\u2013662.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_144_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2015.2439035"},{"key":"e_1_3_2_145_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2017.2649101"},{"key":"e_1_3_2_146_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2016.2575544"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3565267","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3565267","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:43Z","timestamp":1750178263000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3565267"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,25]]},"references-count":145,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,8,31]]}},"alternative-id":["10.1145\/3565267"],"URL":"https:\/\/doi.org\/10.1145\/3565267","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,25]]},"assertion":[{"value":"2022-03-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-09-25","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}