{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:59:44Z","timestamp":1760147984676,"version":"build-2065373602"},"reference-count":48,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2023,3,21]],"date-time":"2023-03-21T00:00:00Z","timestamp":1679356800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China (NSFC), Essential projects","award":["U2033218","61831018"],"award-info":[{"award-number":["U2033218","61831018"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Most popular graph attention networks treat pixels of a feature map as individual nodes, which makes the feature embedding extracted by the graph convolution lack the integrity of the object. Moreover, matching between a template graph and a search graph using only part-level information usually causes tracking errors, especially in occlusion and similarity situations. To address these problems, we propose a novel end-to-end graph attention tracking framework that has high symmetry, combining traditional cross-correlation operations directly. By utilizing cross-correlation operations, we effectively compensate for the dispersion of graph nodes and enhance the representation of features. Additionally, our graph attention fusion model performs both part-to-part matching and global matching, allowing for more accurate information embedding in the template and search regions. Furthermore, we optimize the information embedding between the template and search branches to achieve better single-object tracking results, particularly in occlusion and similarity scenarios. The flexibility of graph nodes and the comprehensiveness of information embedding have brought significant performance improvements in our framework. Extensive experiments on three challenging public datasets (LaSOT, GOT-10k, and VOT2016) show that our tracker outperforms other state-of-the-art trackers.<\/jats:p>","DOI":"10.3390\/sym15030771","type":"journal-article","created":{"date-parts":[[2023,3,22]],"date-time":"2023-03-22T07:46:43Z","timestamp":1679471203000},"page":"771","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Cross-Correlation Fusion Graph Convolution-Based Object Tracking"],"prefix":"10.3390","volume":"15","author":[{"given":"Liuyi","family":"Fan","sequence":"first","affiliation":[{"name":"School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201620, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3089-8208","authenticated-orcid":false,"given":"Wei","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Electrical Engineering & Automation, Jiangsu Normal University, Xuzhou 221116, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoyan","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201620, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"18171","DOI":"10.1007\/s00521-022-07456-2","article-title":"Similarity based person re-identification for multi-object tracking using deep siamese network","volume":"34","author":"Suljagic","year":"2022","journal-title":"Neural Comput. Appl."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"97","DOI":"10.1007\/s00138-022-01349-z","article-title":"Fast re-obj: Real-time object re-identification in rigid scenes","volume":"33","author":"Bayraktar","year":"2022","journal-title":"Mach. Vis. Appl."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Cen, M., and Jung, C. (2018, January 7\u201310). Fully convolutional siamese fusion networks for object tracking. Proceedings of the 2018 25th IEEE International Conference on Image Processing (ICIP), Athens, Greece.","DOI":"10.1109\/ICIP.2018.8451102"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Li, B., Yan, J., Wu, W., Zhu, Z., and Hu, X. (2018, January 18\u201323). High performance visual tracking with siamese region proposal network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00935"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zhu, Z., Wang, Q., Li, B., Wu, W., Yan, J., and Hu, W. (2018, January 8\u201314). Distractor-aware siamese networks for visual object tracking. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01240-3_7"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Li, B., Wu, W., Wang, Q., Zhang, F., Xing, J., and Yan, J. (2019, January 15\u201320). Siamrpn++: Evolution of siamese visual tracking with very deep networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00441"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Xu, Y., Wang, Z., Li, Z., Yuan, Y., and Yu, G. (2020, January 10). Siamfc++: Towards robust and accurate visual tracking with target estimation guidelines. Proceedings of the AAAI Conference on Artificial Intelligence (AIII), New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6944"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Chen, Z., Zhong, B., Li, G., Zhang, S., and Ji, R. (2020, January 14\u201319). Siamese box adaptive network for visual tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00670"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Guo, D., Wang, J., Cui, Y., Wang, Z., and Chen, S. (2020, January 14\u201319). Siamcar: Siamese fully convolutional classification and regression for visual tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00630"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1109\/TNN.2008.2005605","article-title":"The graph neural network model","volume":"20","author":"Scarselli","year":"2008","journal-title":"IEEE Trans. Neural Net."},{"key":"ref_11","first-page":"1","article-title":"Graph convolutional networks: A comprehensive review","volume":"6","author":"Zhang","year":"2019","journal-title":"Comput. Soc. Net."},{"key":"ref_12","first-page":"4","article-title":"Graph attention networks","volume":"1050","author":"Velickovic","year":"2018","journal-title":"Stat"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Gao, J., Zhang, T., and Xu, C. (2019, January 15\u201320). Graph convolutional tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00478"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Guo, D., Shao, Y., Cui, Y., Wang, Z., Zhang, L., and Shen, C. (2021, January 20\u201325). Graph attention tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00942"},{"key":"ref_15","unstructured":"Kristan, J.M.M., Leonardis, A., and Felsberg, M. (2016, January 11\u201314). The visual object tracking vot2016 challenge results. Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1562","DOI":"10.1109\/TPAMI.2019.2957464","article-title":"Got-10k: A large high-diversity benchmark for generic object tracking in the wild","volume":"43","author":"Huang","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell. (Tpami)"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Fan, H., Lin, L., Yang, F., Chu, P., Deng, G., Yu, S., Bai, H., Xu, Y., Liao, C., and Ling, H. (2019, January 15\u201320). Lasot: A high-quality benchmark for large-scale single object tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00552"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Cheng, S., Zhong, B., Li, G., Liu, X., Tang, Z., Li, X., and Wang, J. (2021, January 20\u201325). Learning to filter: Siamese relation network for robust tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00440"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Peng, J., Jiang, Z., Gu, Y., Wu, Y., Wang, Y., Tai, Y., Wang, C., and Lin, W. (2021, January 19\u201320). Siamrcr: Reciprocal classification and regression for visual object tracking. Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), Montr\u00e9al, QC, Canada.","DOI":"10.24963\/ijcai.2021\/132"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Tang, F., and Ling, Q. (2022, January 18\u201324). Ranking-based siamese visual tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00854"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Danelljan, M., Bhat, G., Khan, F.S., and Felsberg, M. (2019, January 15\u201320). Atom: Accurate tracking by overlap maximization. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00479"},{"key":"ref_22","unstructured":"Bhat, G., Danelljan, M., Gool, L.V., and Timofte, R. (November, January 27). Learning discriminative model prediction for tracking. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wang, Y., Kitani, K., and Weng, X. (June, January 30). Joint object detection and multi-object tracking with graph neural networks. Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi\u2019an, China.","DOI":"10.1109\/ICRA48506.2021.9561110"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"13387","DOI":"10.1007\/s00521-022-07368-1","article-title":"Applications of graph convolutional networks in computer vision","volume":"34","author":"Cao","year":"2022","journal-title":"Neural Comput. Appl."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Dai, P., Weng, R., Choi, W., Zhang, C., He, Z., and Ding, W. (2021, January 20\u201325). Learning a proposal classifier for multiple object tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00247"},{"key":"ref_26","unstructured":"Wang, R., Yan, J., and Yang, X. (November, January 27). Learning combinatorial embedding networks for deep graph matching. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Lin, T., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Real, E., Shlens, J., Mazzocchi, S., Pan, X., and Vanhoucke, V. (2017, January 21\u201327). Youtube-boundingboxes: A large high-precision human-annotated data set for object detection in video. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.789"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Zhang, Z., and Peng, H. (2019, January 15\u201320). Deeper and wider siamese networks for real-time visual tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00472"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"247","DOI":"10.1016\/j.isprsjprs.2021.03.005","article-title":"Clnet: Cross-layer convolutional neural network for change detection in optical remote sensing imagery","volume":"175","author":"Zheng","year":"2021","journal-title":"ISPRS J. Photogramm. Remote. Sens. (JPRS)"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zheng, J., Ma, C., Peng, H., and Yang, X. (2021, January 10\u201317). Learning to track objects from unlabeled videos. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01329"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Peng, H., Fu, J., Li, B., and Hu, W. (2020, January 23\u201328). Ocean: Object-aware anchor-free tracking. Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK.","DOI":"10.1007\/978-3-030-58589-1_46"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Lukezic, A., Matas, J., and Kristan, M. (2020, January 13\u201319). D3s-a discriminative single shot segmentation tracker. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00716"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Ko, K., Lee, J.-T., and Kim, C.-S. (2018, January 7\u201310). Pac-net: Pairwise aesthetic comparison network for image aesthetic assessment. Proceedings of the 2018 25th IEEE International Conference on Image Processing (ICIP), Athens, Greece.","DOI":"10.1109\/ICIP.2018.8451621"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Pei, W., Li, X., Wang, H., Zheng, F., and He, Z. (2021, January 10\u201317). Saliency-associated object tracking. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00972"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Yan, B., Peng, H., Wu, K., Wang, D., Fu, J., and Lu, H. (2021, January 20\u201325). Lighttrack: Finding lightweight neural networks for object tracking via one-shot architecture search. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01493"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Yan, B., Peng, H., Fu, J., Wang, D., and Lu, H. (2021, January 11\u201317). Learning spatio-temporal transformer for visual tracking. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01028"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"276","DOI":"10.1016\/j.ins.2022.11.055","article-title":"Tracking vision transformer with class and regression tokens","volume":"619","author":"Nardo","year":"2023","journal-title":"Inf. Sci."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Dai, K., Wang, D., Lu, H., Sun, C., and Li, J. (2019, January 15\u201320). Visual tracking via adaptive spatially-regularized correlation filters. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00480"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Yang, T., Xu, P., Hu, R., Chai, H., and Chan, A.B. (2020, January 13\u201319). Roam: Recurrently optimizing tracking model. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00675"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Shen, Q., Qiao, L., Guo, J., Li, P., Li, X., Li, B., Feng, W., Gan, W., Wu, W., and Ouyang, W. (2022, January 18\u201324). Unsupervised learning of accurate siamese tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00793"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Sio, C.H., Ma, Y., Shuai, H., Chen, J., and Cheng, W. (2020, January 12\u201316). S2siamfc: Self-supervised fully convolutional siamese network for visual tracking. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3413611"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Wang, Q., Zhang, L., Bertinetto, L., Hu, W., and Torr, P.H.S. (2019, January 15\u201320). Fast online object tracking and segmentation: A unifying approach. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00142"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Zhao, W., Deng, M., Cheng, C., and Zhang, D. (2022). Real-time object tracking algorithm based on siamese network. Appl. Sci., 12.","DOI":"10.3390\/app12147338"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Han, W., Dong, X., Khan, F.S., Shao, L., and Shen, J. (2021, January 20\u201325). Learning to fuse asymmetric feature maps in siamese trackers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01630"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Jung, I., You, K., Noh, H., Cho, M., and Han, B. (2020, January 7\u201312). Real-time object tracking via meta-learning: Efficient model adaptation and one-shot channel pruning. Proceedings of the AAAI Conference on Artificial Intelligence (AIII), New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6779"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/15\/3\/771\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:59:51Z","timestamp":1760122791000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/15\/3\/771"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,21]]},"references-count":48,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["sym15030771"],"URL":"https:\/\/doi.org\/10.3390\/sym15030771","relation":{},"ISSN":["2073-8994"],"issn-type":[{"type":"electronic","value":"2073-8994"}],"subject":[],"published":{"date-parts":[[2023,3,21]]}}}