{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T04:54:12Z","timestamp":1783572852978,"version":"3.55.0"},"reference-count":70,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2024,1,22]],"date-time":"2024-01-22T00:00:00Z","timestamp":1705881600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Scientific and Technological Innovation 2030","award":["22022ZD0116407"],"award-info":[{"award-number":["22022ZD0116407"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,5,31]]},"abstract":"<jats:p>\n            There is no doubt that the rational and effective use of visible and thermal infrared image data information to achieve cross-modal complementary fusion is the key to improving the performance of RGB-T salient object detection (SOD). A meticulous analysis of the RGB-T SOD data reveals that it mainly consists of three scenarios in which both modalities (RGB and T) have a significant foreground and only a single modality (RGB or T) is disturbed. However, existing methods are obsessed with pursuing more effective cross-modal fusion based on treating both modalities equally. Obviously, the subjective use of equivalence has two significant limitations. Firstly, it does not allow for practical discrimination of which modality makes the dominant contribution to performance. While both modalities may have visually significant foregrounds, differences in their imaging properties will result in distinct performance contributions. Secondly, in a specific acquisition scenario, a pair of images with two modalities will contribute differently to the final detection performance due to their varying sensitivity to the same background interference. Intelligibly, for the RGB-T saliency detection task, it would be more reasonable to generate exclusive weights for the two modalities and select specific fusion mechanisms based on different weight configurations to perform cross-modal complementary integration. Consequently, we propose a weighted guided optional fusion network (WGOFNet) for RGB-T SOD. Specifically, a feature refinement module is first used to perform an initial refinement of the extracted multilevel features. Subsequently, a weight generation module (WGM) will generate exclusive network performance contribution weights for each of the two modalities, and an optional fusion module (OFM) will rely on this weight to perform particular integration of cross-modal information. Simple cross-level fusion is finally utilized to obtain the final saliency prediction map. Comprehensive experiments on three publicly available benchmark datasets demonstrate the proposed WGOFNet achieves superior performance compared with the state-of-the-art RGB-T SOD methods. The source code is available at:\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/WJ-CV\/WGOFNet\">https:\/\/github.com\/WJ-CV\/WGOFNet<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3624984","type":"journal-article","created":{"date-parts":[[2023,10,13]],"date-time":"2023-10-13T15:26:17Z","timestamp":1697210777000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["Weighted Guided Optional Fusion Network for RGB-T Salient Object Detection"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7266-8179","authenticated-orcid":false,"given":"Jie","family":"Wang","sequence":"first","affiliation":[{"name":"College of Intelligence and Computing, Tianjin University, Tianjin, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6671-6720","authenticated-orcid":false,"given":"Guoqiang","family":"Li","sequence":"additional","affiliation":[{"name":"Institutes of Science and Development, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-0433-3397","authenticated-orcid":false,"given":"Jie","family":"Shi","sequence":"additional","affiliation":[{"name":"North Automatic Control Technology Institute, Taiyuan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7504-3457","authenticated-orcid":false,"given":"Jinwen","family":"Xi","sequence":"additional","affiliation":[{"name":"Zhongguancun Laboratory, Beijing, P.R.China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,1,22]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"crossref","first-page":"1597","DOI":"10.1109\/CVPR.2009.5206596","volume-title":"Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition","author":"Achanta Radhakrishna","year":"2009","unstructured":"Radhakrishna Achanta, Sheila Hemami, Francisco Estrada, and Sabine Susstrunk. 2009. Frequency-tuned salient region detection. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1597\u20131604. DOI:10.1109\/CVPR.2009.5206596"},{"key":"e_1_3_1_3_2","article-title":"Visible and thermal images fusion architecture for few-shot semantic segmentation","volume":"80","author":"Bao Yanqi","year":"2021","unstructured":"Yanqi Bao, Kechen Song, Jie Wang, Liming Huang, Hongwen Dong, and Yunhui Yan. 2021. Visible and thermal images fusion architecture for few-shot semantic segmentation. Journal of Visual Communication and Image Representation 80 (2021), 103306.","journal-title":"Journal of Visual Communication and Image Representation"},{"issue":"7","key":"e_1_3_1_4_2","doi-asserted-by":"crossref","first-page":"3156","DOI":"10.1109\/TIP.2017.2670143","article-title":"Video saliency detection via spatial-temporal fusion and low-rank coherency diffusion","volume":"26","author":"Chen Chenglizhao","year":"2017","unstructured":"Chenglizhao Chen, Shuai Li, Yongguang Wang, Hong Qin, and Aimin Hao. 2017. Video saliency detection via spatial-temporal fusion and low-rank coherency diffusion. IEEE Transactions on Image Processing 26, 7 (2017), 3156\u20133170.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_5_2","doi-asserted-by":"crossref","first-page":"3995","DOI":"10.1109\/TIP.2021.3068644","article-title":"Exploring rich and efficient spatial temporal interactions for real-time video salient object detection","volume":"30","author":"Chen Chenglizhao","year":"2021","unstructured":"Chenglizhao Chen, Guotao Wang, Chong Peng, Yuming Fang, Dingwen Zhang, and Hong Qin. 2021. Exploring rich and efficient spatial temporal interactions for real-time video salient object detection. IEEE Transactions on Image Processing 30 (2021), 3995\u20134007.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","first-page":"1090","DOI":"10.1109\/TIP.2019.2934350","article-title":"Improved robust video saliency detection based on long-term spatial-temporal information","volume":"29","author":"Chen Chenglizhao","year":"2019","unstructured":"Chenglizhao Chen, Guotao Wang, Chong Peng, Xiaowei Zhang, and Hong Qin. 2019. Improved robust video saliency detection based on long-term spatial-temporal information. IEEE Transactions on Image Processing 29 (2019), 1090\u20131100.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_7_2","doi-asserted-by":"crossref","first-page":"2350","DOI":"10.1109\/TIP.2021.3052069","article-title":"Depth-quality-aware salient object detection","volume":"30","author":"Chen Chenglizhao","year":"2021","unstructured":"Chenglizhao Chen, Jipeng Wei, Chong Peng, and Hong Qin. 2021. Depth-quality-aware salient object detection. IEEE Transactions on Image Processing 30 (2021), 2350\u20132363.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","first-page":"4296","DOI":"10.1109\/TIP.2020.2968250","article-title":"Improved saliency detection in RGB-D images using two-phase depth estimation and selective deep fusion","volume":"29","author":"Chen Chenglizhao","year":"2020","unstructured":"Chenglizhao Chen, Jipeng Wei, Chong Peng, Weizhong Zhang, and Hong Qin. 2020. Improved saliency detection in RGB-D images using two-phase depth estimation and selective deep fusion. IEEE Transactions on Image Processing 29 (2020), 4296\u20134307.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3166914"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2699184"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","first-page":"520","DOI":"10.1007\/978-3-030-58598-3_31","volume-title":"Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference","author":"Chen Shuhan","year":"2020","unstructured":"Shuhan Chen and Yun Fu. 2020. Progressively guided alternate refinement network for RGB-D salient object detection. In Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference. Springer, 520\u2013538. DOI:10.1007\/978-3-030-58598-3_31"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3028289"},{"key":"e_1_3_1_13_2","first-page":"3286","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Cheng Ming-Ming","year":"2014","unstructured":"Ming-Ming Cheng, Ziming Zhang, Wen-Yan Lin, and Philip Torr. 2014. BING: Binarized normed gradients for objectness estimation at 300fps. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3286\u20133293."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3216476"},{"key":"e_1_3_1_15_2","first-page":"4548","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Fan Deng-Ping","year":"2017","unstructured":"Deng-Ping Fan, Ming-Ming Cheng, Yun Liu, Tao Li, and Ali Borji. 2017. Structure-measure: A new way to evaluate foreground maps. In Proceedings of the IEEE International Conference on Computer Vision. 4548\u20134557."},{"key":"e_1_3_1_16_2","unstructured":"Deng-Ping Fan Cheng Gong Yang Cao Bo Ren Ming-Ming Cheng and Ali Borji. 2018. Enhanced-alignment measure for binary foreground map evaluation. arXiv:1805.10421. Retrieved from https:\/\/arxiv.org\/abs\/1805.10421"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.108666"},{"key":"e_1_3_1_18_2","first-page":"3052","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Fu Keren","year":"2020","unstructured":"Keren Fu, Deng-Ping Fan, Ge-Peng Ji, and Qijun Zhao. 2020. JL-DCF: Joint learning and densely-cooperative fusion framework for RGB-D salient object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3052\u20133062."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3082939"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2015.2389616"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2021.3102524"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2020.3020735"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3069812"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3069297"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3102268"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIM.2022.3185323s"},{"key":"e_1_3_1_27_2","first-page":"9471","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ji Wei","year":"2021","unstructured":"Wei Ji, Jingjing Li, Shuang Yu, Miao Zhang, Yongri Piao, Shunyu Yao, Qi Bi, Kai Ma, Yefeng Zheng, Huchuan Lu, et\u00a0al. 2021. Calibrated RGB-D salient object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 9471\u20139481."},{"key":"e_1_3_1_28_2","first-page":"3496","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Jia Xinyu","year":"2021","unstructured":"Xinyu Jia, Chuang Zhu, Minzhen Li, Wenqi Tang, and Wenli Zhou. 2021. LLVIP: A visible-infrared paired dataset for low-light vision. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 3496\u20133504."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3062689"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.03.029"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3184840"},{"key":"e_1_3_1_32_2","first-page":"5802","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu Jinyuan","year":"2022","unstructured":"Jinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu, Risheng Liu, Wei Zhong, and Zhongxuan Luo. 2022. Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 5802\u20135811."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3140168"},{"key":"e_1_3_1_34_2","first-page":"4722","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Liu Nian","year":"2021","unstructured":"Nian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao, and Junwei Han. 2021. Visual saliency transformer. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4722\u20134732."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3127149"},{"key":"e_1_3_1_36_2","first-page":"4481","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Liu Zhengyi","year":"2021","unstructured":"Zhengyi Liu, Yuan Wang, Zhengzheng Tu, Yun Xiao, and Bin Tang. 2021. TriTransNet: RGB-D salient object detection with a triplet transformer embedding network. In Proceedings of the 29th ACM International Conference on Multimedia. 4481\u20134490. DOI:10.1145\/3474085.3475601"},{"key":"e_1_3_1_37_2","first-page":"248","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Margolin Ran","year":"2014","unstructured":"Ran Margolin, Lihi Zelnik-Manor, and Ayellet Tal. 2014. How to evaluate foreground maps?. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 248\u2013255."},{"key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"892","DOI":"10.1109\/TIP.2023.3234702","article-title":"CAVER: Cross-modal view-mixed transformer for bi-modal salient object detection","volume":"32","author":"Pang Youwei","year":"2023","unstructured":"Youwei Pang, Xiaoqi Zhao, Lihe Zhang, and Huchuan Lu. 2023. CAVER: Cross-modal view-mixed transformer for bi-modal salient object detection. IEEE Transactions on Image Processing 32 (2023), 892\u2013904.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","first-page":"733","DOI":"10.1109\/CVPR.2012.6247743","volume-title":"Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition","author":"Perazzi Federico","year":"2012","unstructured":"Federico Perazzi, Philipp Kr\u00e4henb\u00fchl, Yael Pritch, and Alexander Hornung. 2012. Saliency filters: Contrast based filtering for salient region detection. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 733\u2013740. DOI:10.1109\/CVPR.2012.6247743"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2006.04.010"},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TIM.2023.3236346","article-title":"A potential vision-based measurements technology: Information flow fusion detection method using RGB-thermal infrared images","volume":"72","author":"Song Kechen","year":"2023","unstructured":"Kechen Song, Yanqi Bao, Han Wang, Liming Huang, and Yunhui Yan. 2023. A potential vision-based measurements technology: Information flow fusion detection method using RGB-thermal infrared images. IEEE Transactions on Instrumentation and Measurement 72 (2023), 1\u201313.","journal-title":"IEEE Transactions on Instrumentation and Measurement"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMECH.2022.3215909"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSP.2022.3194843"},{"key":"e_1_3_1_44_2","first-page":"1407","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Sun Peng","year":"2021","unstructured":"Peng Sun, Wenhu Zhang, Huanyu Wang, Songyuan Li, and Xi Li. 2021. Deep RGB-D saliency detection with depth-sensitive attention and automatic multi-modal fusion. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 1407\u20131417."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3087412"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3176540"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3171688"},{"key":"e_1_3_1_48_2","first-page":"141","volume-title":"Proceedings of the 2019 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR)","author":"Tu Zhengzheng","year":"2019","unstructured":"Zhengzheng Tu, Tian Xia, Chenglong Li, Yijuan Lu, and Jin Tang. 2019. M3S-NIR: Multi-modal multi-scale noise-insensitive ranking for RGB-T saliency detection. In Proceedings of the 2019 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR). IEEE, 141\u2013146. DOI:10.1109\/MIPR.2019.00032"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2924578"},{"key":"e_1_3_1_50_2","first-page":"15119","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Guotao","year":"2021","unstructured":"Guotao Wang, Chenglizhao Chen, Deng-Ping Fan, Aimin Hao, and Hong Qin. 2021. From semantic categories to fixations: A novel weakly-supervised visual-auditory saliency detection approach. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 15119\u201315128."},{"key":"e_1_3_1_51_2","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1007\/978-981-13-1702-6_36","volume-title":"Proceedings of the Image and Graphics Technologies and Applications: 13th Conference on Image and Graphics Technologies and Applications, IGTA 2018","author":"Wang Guizhao","year":"2018","unstructured":"Guizhao Wang, Chenglong Li, Yunpeng Ma, Aihua Zheng, Jin Tang, and Bin Luo. 2018. RGB-T saliency detection benchmark: Dataset, baselines, analysis and a novel approach. In Proceedings of the Image and Graphics Technologies and Applications: 13th Conference on Image and Graphics Technologies and Applications, IGTA 2018. Springer, 359\u2013369. DOI:10.1007\/978-981-13-1702-6_36"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3099120"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2022.105162"},{"key":"e_1_3_1_54_2","first-page":"568","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Wang Wenhai","year":"2021","unstructured":"Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2021. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 568\u2013578."},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3123548"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3134684"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3164550"},{"issue":"2","key":"e_1_3_1_58_2","doi-asserted-by":"crossref","first-page":"535","DOI":"10.1049\/ipr2.12369","article-title":"LBP-based progressive feature aggregation network for low-light image enhancement","volume":"16","author":"Yu Nana","year":"2022","unstructured":"Nana Yu, Jinjiang Li, and Zhen Hua. 2022. LBP-based progressive feature aggregation network for low-light image enhancement. IET Image Processing 16, 2 (2022), 535\u2013553.","journal-title":"IET Image Processing"},{"issue":"4","key":"e_1_3_1_59_2","first-page":"1251","article-title":"Fla-net: Multi-stage modular network for low-light image enhancement","volume":"39","author":"Yu Nana","year":"2023","unstructured":"Nana Yu, Jinjiang Li, and Zhen Hua. 2023. Fla-net: Multi-stage modular network for low-light image enhancement. The Visual Computer 39, 4 (2023), 1251\u20131270.","journal-title":"The Visual Computer"},{"key":"e_1_3_1_60_2","first-page":"4338","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhang Jing","year":"2021","unstructured":"Jing Zhang, Deng-Ping Fan, Yuchao Dai, Xin Yu, Yiran Zhong, Nick Barnes, and Ling Shao. 2021. RGB-D saliency detection via cascaded mutual information minimization. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4338\u20134347."},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2019.107130"},{"key":"e_1_3_1_62_2","first-page":"8886","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang Pengyu","year":"2022","unstructured":"Pengyu Zhang, Jie Zhao, Dong Wang, Huchuan Lu, and Xiang Ruan. 2022. Visible-thermal UAV tracking: A large-scale benchmark and new baseline. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 8886\u20138895."},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2959253"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2020.3014663"},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","first-page":"1745","DOI":"10.1145\/3394171.3413855","volume-title":"Proceedings of the 28th ACM International Conference on Multimedia","author":"Zhao Jiawei","year":"2020","unstructured":"Jiawei Zhao, Yifan Zhao, Jia Li, and Xiaowu Chen. 2020. Is depth really necessary for salient object detection?. In Proceedings of the 28th ACM International Conference on Multimedia. 1745\u20131754. DOI:10.1145\/3394171.3413855"},{"key":"e_1_3_1_66_2","article-title":"Position-Aware Relation Learning for RGB-Thermal Salient Object Detection","author":"Zhou Heng","year":"2023","unstructured":"Heng Zhou, Chunna Tian, Zhenxi Zhang, Chengyang Li, Yuxuan Ding, Yongqiang Xie, and Zhongbo Li. 2023. Position-Aware Relation Learning for RGB-Thermal Salient Object Detection. IEEE Transactions on Image Processing (2023).","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-020-0199-z"},{"key":"e_1_3_1_68_2","first-page":"4681","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhou Tao","year":"2021","unstructured":"Tao Zhou, Huazhu Fu, Geng Chen, Yi Zhou, Deng-Ping Fan, and Ling Shao. 2021. Specificity-preserving RGB-D saliency detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4681\u20134691."},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3077058"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/TETCI.2021.3118043"},{"key":"e_1_3_1_71_2","article-title":"Saliency Prototype for RGB-D and RGB-T Salient Object Detection","author":"Wang Yahong Han Zihao Zhang, Jie","year":"2023","unstructured":"Yahong Han Zihao Zhang, Jie Wang. 2023. Saliency Prototype for RGB-D and RGB-T Salient Object Detection. ACM International Conference on Multimedia (2023).","journal-title":"ACM International Conference on Multimedia"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3624984","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3624984","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:35:45Z","timestamp":1750178145000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3624984"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,22]]},"references-count":70,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,5,31]]}},"alternative-id":["10.1145\/3624984"],"URL":"https:\/\/doi.org\/10.1145\/3624984","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,22]]},"assertion":[{"value":"2023-05-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-10","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}