{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T19:40:06Z","timestamp":1783107606499,"version":"3.54.6"},"reference-count":44,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2024,3,11]],"date-time":"2024-03-11T00:00:00Z","timestamp":1710115200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,3,11]],"date-time":"2024-03-11T00:00:00Z","timestamp":1710115200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Key Science and Technology Project of Henan Province","award":["201300210400"],"award-info":[{"award-number":["201300210400"]}]},{"DOI":"10.13039\/501100017700","name":"Henan Province Science and Technology Research Project","doi-asserted-by":"crossref","award":["232102210031"],"award-info":[{"award-number":["232102210031"]}],"id":[{"id":"10.13039\/501100017700","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Process Lett"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The Siamese network-based tracker calculates object templates and search images independently, and the template features are not updated online when performing object tracking. Adapting to interference scenarios with performance-guaranteed tracking accuracy when background clutter, illumination variation or partial occlusion occurs in the search area is a challenging task. To effectively address the issue with the abovementioned interference and to improve location accuracy, this paper devises a Siamese residual attentional aggregation network framework for self-adaptive feature implicit updating. First, SiamRAAN introduces Self-RAAN into the backbone network by applying residual self-attention to extract effective objective features. Then, we introduce Cross-RAAN to update the template features online by focusing on the high-relevance parts in the feature extraction process of both the object template and search image. Finally, a multilevel feature fusion module is introduced to fuse the RAAN-enhanced feature information and improve the network\u2019s ability to perceive key features. Extensive experiments conducted on benchmark datasets (GOT-10K, LaSOT, OTB-50, OTB-100 and UAV123) demonstrated that our SiamRAAN delivers excellent performance and runs at 51 FPS in various challenging object tracking tasks. Code is available at <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/MallowYi\/SiamRAAN\">https:\/\/github.com\/MallowYi\/SiamRAAN<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/s11063-024-11556-6","type":"journal-article","created":{"date-parts":[[2024,3,11]],"date-time":"2024-03-11T09:01:32Z","timestamp":1710147692000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["SiamRAAN: Siamese Residual Attentional Aggregation Network for Visual Object Tracking"],"prefix":"10.1007","volume":"56","author":[{"given":"Zhiyi","family":"Xin","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Junyang","family":"Yu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xin","family":"He","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yalin","family":"Song","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Han","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,3,11]]},"reference":[{"key":"11556_CR1","doi-asserted-by":"crossref","unstructured":"Held D, Thrun S, Savarese S (2016) Learning to track at 100 fps with deep regression networks. In: Computer vision\u2013ECCV 2016: 14th European conference, Amsterdam, The Netherlands, 11\u201314 Oct 2016, Proceedings, Part I 14, pp 749\u2013765. Springer","DOI":"10.1007\/978-3-319-46448-0_45"},{"issue":"3","key":"11556_CR2","doi-asserted-by":"publisher","first-page":"583","DOI":"10.1109\/TPAMI.2014.2345390","volume":"37","author":"JF Henriques","year":"2014","unstructured":"Henriques JF, Caseiro R, Martins P, Batista J (2014) High-speed tracking with kernelized correlation filters. IEEE Trans Pattern Anal Mach Intell 37(3):583\u2013596","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11556_CR3","doi-asserted-by":"crossref","unstructured":"Guo D, Wang J, Cui Y, Wang Z, Chen S (2020) Siamcar: siamese fully convolutional classification and regression for visual tracking. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 6269\u20136277","DOI":"10.1109\/CVPR42600.2020.00630"},{"key":"11556_CR4","doi-asserted-by":"crossref","unstructured":"Wang F, Cao P, Wang X, He B, Sun F (2023) Siamadt: siamese attention and deformable features fusion network for visual object tracking. Neural Process Lett, pp 1\u201318","DOI":"10.21203\/rs.3.rs-2190588\/v1"},{"issue":"2","key":"11556_CR5","doi-asserted-by":"publisher","first-page":"1029","DOI":"10.1007\/s11063-022-10924-4","volume":"55","author":"Y Wu","year":"2023","unstructured":"Wu Y, Cai C, Yeo CK (2023) Siamese centerness prediction network for real-time visual object tracking. Neural Process Lett 55(2):1029\u20131044","journal-title":"Neural Process Lett"},{"issue":"5","key":"11556_CR6","doi-asserted-by":"publisher","first-page":"2019","DOI":"10.1109\/TIP.2014.2311377","volume":"23","author":"J Yu","year":"2014","unstructured":"Yu J, Rui Y, Tao D (2014) Click prediction for web image reranking using multimodal sparse coding. IEEE Trans Image Process 23(5):2019\u20132032","journal-title":"IEEE Trans Image Process"},{"issue":"2","key":"11556_CR7","doi-asserted-by":"publisher","first-page":"563","DOI":"10.1109\/TPAMI.2019.2932058","volume":"44","author":"J Yu","year":"2019","unstructured":"Yu J, Tan M, Zhang H, Rui Y, Tao D (2019) Hierarchical deep click feature prediction for fine-grained image recognition. IEEE Trans Pattern Anal Mach Intell 44(2):563\u2013578","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11556_CR8","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.107952","volume":"116","author":"J Zhang","year":"2021","unstructured":"Zhang J, Cao Y, Wu Q (2021) Vector of locally and adaptively aggregated descriptors for image feature representation. Pattern Recogn 116:107952","journal-title":"Pattern Recogn"},{"issue":"5","key":"11556_CR9","doi-asserted-by":"publisher","first-page":"3117","DOI":"10.1002\/int.22814","volume":"37","author":"J Zhang","year":"2022","unstructured":"Zhang J, Yang J, Yu J, Fan J (2022) Semisupervised image classification by mutual learning of multiple self-supervised models. Int J Intell Syst 37(5):3117\u20133141","journal-title":"Int J Intell Syst"},{"key":"11556_CR10","doi-asserted-by":"crossref","unstructured":"Marvasti-Zadeh SM, Cheng L, Ghanei-Yakhdan H, Kasaei S (2021) Deep learning for visual tracking: A comprehensive survey. IEEE Trans Intell Transp Syst","DOI":"10.1109\/TITS.2020.3046478"},{"key":"11556_CR11","doi-asserted-by":"publisher","first-page":"39","DOI":"10.1016\/j.neucom.2020.01.085","volume":"396","author":"X Wu","year":"2020","unstructured":"Wu X, Sahoo D, Hoi SC (2020) Recent advances in deep learning for object detection. Neurocomputing 396:39\u201364","journal-title":"Neurocomputing"},{"key":"11556_CR12","doi-asserted-by":"crossref","unstructured":"Bertinetto L, Valmadre J, Henriques JF, Vedaldi A, Torr PH (2016) Fully-convolutional siamese networks for object tracking. In: Computer vision\u2013ECCV 2016 workshops: Amsterdam, The Netherlands, 8\u201310 and 15\u201316, Oct 2016, Proceedings, Part II 14, pp 850\u2013865. Springer","DOI":"10.1007\/978-3-319-48881-3_56"},{"key":"11556_CR13","doi-asserted-by":"publisher","first-page":"2114","DOI":"10.1109\/TMM.2020.3008028","volume":"23","author":"Q Liu","year":"2020","unstructured":"Liu Q, Li X, He Z, Fan N, Yuan D, Wang H (2020) Learning deep multi-level similarity for thermal infrared object tracking. IEEE Trans Multimedia 23:2114\u20132126","journal-title":"IEEE Trans Multimedia"},{"key":"11556_CR14","doi-asserted-by":"crossref","unstructured":"Li B, Yan J, Wu W, Zhu Z, Hu X (2018) High performance visual tracking with siamese region proposal network. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 8971\u20138980","DOI":"10.1109\/CVPR.2018.00935"},{"key":"11556_CR15","doi-asserted-by":"crossref","unstructured":"Zhu Z, Wang Q, Li B, Wu W, Yan J, Hu W (2018) Distractor-aware siamese networks for visual object tracking. In: Proceedings of the European conference on computer vision (ECCV), pp 101\u2013117","DOI":"10.1007\/978-3-030-01240-3_7"},{"key":"11556_CR16","doi-asserted-by":"crossref","unstructured":"Li B, Wu W, Wang Q, Zhang F, Xing J, Yan J (2019) Siamrpn++: evolution of siamese visual tracking with very deep networks. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 4282\u20134291","DOI":"10.1109\/CVPR.2019.00441"},{"key":"11556_CR17","doi-asserted-by":"crossref","unstructured":"Fan H, Ling H (2019) Siamese cascaded region proposal networks for real-time visual tracking. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 7952\u20137961","DOI":"10.1109\/CVPR.2019.00814"},{"key":"11556_CR18","doi-asserted-by":"crossref","unstructured":"Guo Q, Feng W, Zhou C, Huang R, Wan L, Wang S (2017) Learning dynamic siamese network for visual object tracking. In: Proceedings of the IEEE international conference on computer vision, pp 1763\u20131771","DOI":"10.1109\/ICCV.2017.196"},{"key":"11556_CR19","doi-asserted-by":"crossref","unstructured":"Yang T, Chan AB (2018) Learning dynamic memory networks for object tracking. In: Proceedings of the European conference on computer vision (ECCV), pp 152\u2013167","DOI":"10.1007\/978-3-030-01240-3_10"},{"key":"11556_CR20","doi-asserted-by":"crossref","unstructured":"Wang Q, Teng Z, Xing J, Gao J, Hu W, Maybank S (2018) Learning attentions: residual attentional siamese network for high performance online visual tracking. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 4854\u20134863","DOI":"10.1109\/CVPR.2018.00510"},{"key":"11556_CR21","doi-asserted-by":"publisher","DOI":"10.1016\/j.dsp.2021.103269","volume":"120","author":"J Yu","year":"2022","unstructured":"Yu J, Zuo M, Dong L, Zhang H, He X (2022) The multi-level classification and regression network for visual tracking via residual channel attention. Digital Signal Process 120:103269","journal-title":"Digital Signal Process"},{"key":"11556_CR22","doi-asserted-by":"crossref","unstructured":"Tao R, Gavves E, Smeulders AW (2016) Siamese instance search for tracking. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1420\u20131429","DOI":"10.1109\/CVPR.2016.158"},{"key":"11556_CR23","unstructured":"Ren S, He K, Girshick R, Sun J (2015) Faster r-cnn: Towards real-time object detection with region proposal networks. Adv Neural Inf Process Syst 28"},{"key":"11556_CR24","doi-asserted-by":"crossref","unstructured":"Zhang Z, Peng H (2019) Deeper and wider siamese networks for real-time visual tracking. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 4591\u20134600","DOI":"10.1109\/CVPR.2019.00472"},{"key":"11556_CR25","doi-asserted-by":"crossref","unstructured":"Guo D, Shao Y, Cui Y, Wang Z, Zhang L, Shen C (2021) Graph attention tracking. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 9543\u20139552","DOI":"10.1109\/CVPR46437.2021.00942"},{"issue":"6","key":"11556_CR26","doi-asserted-by":"publisher","DOI":"10.1117\/1.JEI.31.6.063022","volume":"31","author":"Z Zhao","year":"2022","unstructured":"Zhao Z, Zuo M, Yu J, He X, Song Y, Zhai R (2022) Siamese network based on global and local feature matching for object tracking. J Electronic Imag 31(6):063022","journal-title":"J Electronic Imag"},{"key":"11556_CR27","doi-asserted-by":"crossref","unstructured":"Du F, Liu P, Zhao W, Tang X (2020) Correlation-guided attention for corner detection based visual tracking. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 6836\u20136845","DOI":"10.1109\/CVPR42600.2020.00687"},{"key":"11556_CR28","doi-asserted-by":"crossref","unstructured":"Yu Y, Xiong Y, Huang W, Scott MR (2020) Deformable siamese attention networks for visual object tracking. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 6728\u20136737","DOI":"10.1109\/CVPR42600.2020.00676"},{"key":"11556_CR29","doi-asserted-by":"crossref","unstructured":"Yang S, Chen H, Xu F, Li Y, Yuan J (2022) High-performance uavs visual tracking based on siamese network. Visual Comput, pp 1\u201317","DOI":"10.1007\/s00371-021-02271-7"},{"issue":"7","key":"11556_CR30","doi-asserted-by":"publisher","first-page":"2555","DOI":"10.1007\/s00371-021-02131-4","volume":"38","author":"C Guo","year":"2022","unstructured":"Guo C, Yang D, Li C, Song P (2022) Dual siamese network for rgbt tracking via fusing predicted position maps. Visual Comput 38(7):2555\u20132567","journal-title":"Visual Comput"},{"key":"11556_CR31","doi-asserted-by":"crossref","unstructured":"Pang H, Han L, Liu C, Ma R (2023) Siamese object tracking based on multi-frequency enhancement feature. Visual Comput, pp 1\u201311","DOI":"10.1007\/s00371-023-02779-0"},{"key":"11556_CR32","doi-asserted-by":"crossref","unstructured":"Woo S, Park J, Lee J-Y, Kweon IS (2018) Cbam: Convolutional block attention module. In: Proceedings of the European conference on computer vision (ECCV), pp 3\u201319","DOI":"10.1007\/978-3-030-01234-2_1"},{"issue":"11","key":"11556_CR33","doi-asserted-by":"publisher","first-page":"2709","DOI":"10.1109\/TPAMI.2018.2865311","volume":"41","author":"C Ma","year":"2018","unstructured":"Ma C, Huang J-B, Yang X, Yang M-H (2018) Robust visual tracking via hierarchical convolutional features. IEEE Trans Pattern Anal Mach Intell 41(11):2709\u20132723","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11556_CR34","doi-asserted-by":"crossref","unstructured":"Wang G, Luo C, Xiong Z, Zeng W (2019) Spm-tracker: Series-parallel matching for real-time visual object tracking. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 3643\u20133652","DOI":"10.1109\/CVPR.2019.00376"},{"key":"11556_CR35","doi-asserted-by":"crossref","unstructured":"Ma C, Huang J-B, Yang X, Yang M-H (2015) Hierarchical convolutional features for visual tracking. In: Proceedings of the IEEE international conference on computer vision, pp 3074\u20133082","DOI":"10.1109\/ICCV.2015.352"},{"key":"11556_CR36","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky O, Deng J, Su H, Krause J, Satheesh S, Ma S, Huang Z, Karpathy A, Khosla A, Bernstein M et al (2015) Imagenet large scale visual recognition challenge. Int J Comput Vis 115:211\u2013252","journal-title":"Int J Comput Vis"},{"key":"11556_CR37","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll\u00e1r P, Zitnick CL (2014) Microsoft coco: Common objects in context. In: Computer Vision\u2013ECCV 2014: 13th European conference, Zurich, Switzerland, 6\u201312 Sept 2014, Proceedings, Part V 13, pp 740\u2013755. Springer","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"11556_CR38","doi-asserted-by":"crossref","unstructured":"Real E, Shlens J, Mazzocchi S, Pan X, Vanhoucke V (2017) Youtube-boundingboxes: a large high-precision human-annotated data set for object detection in video. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 5296\u20135305","DOI":"10.1109\/CVPR.2017.789"},{"issue":"5","key":"11556_CR39","doi-asserted-by":"publisher","first-page":"1562","DOI":"10.1109\/TPAMI.2019.2957464","volume":"43","author":"L Huang","year":"2019","unstructured":"Huang L, Zhao X, Huang K (2019) Got-10k: A large high-diversity benchmark for generic object tracking in the wild. IEEE Trans Pattern Anal Mach Intell 43(5):1562\u20131577","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11556_CR40","doi-asserted-by":"crossref","unstructured":"Fan H, Lin L, Yang F, Chu P, Deng G, Yu S, Bai H, Xu Y, Liao C, Ling H (2019) Lasot: a high-quality benchmark for large-scale single object tracking. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 5374\u20135383","DOI":"10.1109\/CVPR.2019.00552"},{"key":"11556_CR41","doi-asserted-by":"crossref","unstructured":"Wu Y, Lim J, Yang M-H (2013) Online object tracking: A benchmark. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 2411\u20132418","DOI":"10.1109\/CVPR.2013.312"},{"key":"11556_CR42","doi-asserted-by":"crossref","unstructured":"Wu Y, Lim J, Yang M (2015) Object tracking benchmark. IEEE Trans Pattern Anal Mach Intell","DOI":"10.1109\/TPAMI.2014.2388226"},{"key":"11556_CR43","doi-asserted-by":"crossref","unstructured":"Mueller M, Smith N, Ghanem B (2016) A benchmark and simulator for uav tracking. In: Computer vision\u2013ECCV 2016: 14th European conference, Amsterdam, The Netherlands, 11\u201314 Oct 2016, Proceedings, Part I 14, pp 445\u2013461. Springer","DOI":"10.1007\/978-3-319-46448-0_27"},{"key":"11556_CR44","doi-asserted-by":"crossref","unstructured":"Li X, Ma C, Wu B, He Z, Yang M-H (2019) Target-aware deep tracking. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 1369\u20131378","DOI":"10.1109\/CVPR.2019.00146"}],"container-title":["Neural Processing Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-024-11556-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11063-024-11556-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-024-11556-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,16]],"date-time":"2024-05-16T20:35:52Z","timestamp":1715891752000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11063-024-11556-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,11]]},"references-count":44,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2024,4]]}},"alternative-id":["11556"],"URL":"https:\/\/doi.org\/10.1007\/s11063-024-11556-6","relation":{},"ISSN":["1573-773X"],"issn-type":[{"value":"1573-773X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,11]]},"assertion":[{"value":"11 February 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 March 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"98"}}