{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T13:25:47Z","timestamp":1784640347629,"version":"3.55.0"},"reference-count":97,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T00:00:00Z","timestamp":1775865600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T00:00:00Z","timestamp":1775865600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach. Intell. Res."],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Improving the tolerance for accumulated errors in object localization and enhancing the stability for estimating matching relations are key issues in robust Siamese network-based single object tracking in real-time. In this paper, we propose an object-aware anchor-free network for object tracking. Different from refining the reference anchors, the position and scale of the object are predicted in an anchor-free way. Since each position within a ground-truth box is trained, inaccurate predictions of the object can be rectified during tracking. An irregular sampling-based alignment module is designed to extract object-aware features from the predicted bounding box. The object-aware features are used to improve the classification of the predicted bounding box into the object or background. In order to improve the stability of matching-relation learning for object tracking, we upgrade the Siamese network by using automated search of the matching network based on binary channel operations. Not relying on explicit similarity calculation, by combining relation-operators, different matching networks are searched for the classification and regression tasks for tracking. The adaptability of relation-learning to different tasks is enhanced. Experimental results on many tracking benchmarks show the effectiveness of the proposed trackers.<\/jats:p>","DOI":"10.1007\/s11633-026-1634-0","type":"journal-article","created":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T13:50:22Z","timestamp":1775915422000},"page":"565-592","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Ocean: Object-aware Anchor-free Tracking with Matching-relation Learning"],"prefix":"10.1007","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9237-8825","authenticated-orcid":false,"given":"Weiming","family":"Hu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhipeng","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bing","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Houwen","family":"Peng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stephen","family":"Maybank","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,4,11]]},"reference":[{"key":"1634_CR1","doi-asserted-by":"publisher","first-page":"367","DOI":"10.1007\/978-3-031-73347-5_21","volume-title":"Proceedings of the 18th European Conference on Computer Vision","author":"N Tumanyan","year":"2025","unstructured":"N. Tumanyan, A. Singer, S. Bagon, T. Dekel. DINO-Tracker: Taming DINO for self-supervised point tracking in a single video. In Proceedings of the 18th European Conference on Computer Vision, Milan, Italy, pp. 367\u2013385, 2025. DOI: https:\/\/doi.org\/10.1007\/978-3-031-73347-5_21."},{"key":"1634_CR2","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Y Cui","year":"2023","unstructured":"Y. Cui, T. Song, G. Wu, L. Wang. MixFormerV2: Efficient fully transformer tracking. In Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, USA, Article number 2561, 2023."},{"issue":"7","key":"1634_CR3","doi-asserted-by":"publisher","first-page":"3918","DOI":"10.1007\/s11263-024-02327-w","volume":"133","author":"J Gao","year":"2025","unstructured":"J. Gao, S. Lin, S. Wang, Y. Kou, Z. Li, L. Li, C. Zhang, X. Zhang, Y. Wang, W. Hu. An experimental study on exploring strong lightweight vision transformers via masked image modeling pre-training. International Journal of Computer Vision, vol. 133, no. 7, pp. 3918\u20133950, 2025. DOI: https:\/\/doi.org\/10.1007\/s11263-024-02327-w.","journal-title":"International Journal of Computer Vision"},{"key":"1634_CR4","doi-asserted-by":"publisher","first-page":"6694","DOI":"10.1109\/WACV57701.2024.00657","volume-title":"Proceedings of IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"G Y Gopal","year":"2024","unstructured":"G. Y. Gopal, M. A. Amer. Separable self and mixed attention transformers for efficient object tracking. In Proceedings of IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, USA, pp. 6694\u20136703, 2024. DOI: https:\/\/doi.org\/10.1109\/WACV57701.2024.00657."},{"key":"1634_CR5","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Y Li","year":"2024","unstructured":"Y. Li, M. Liu, Y. Wu, X. Wang, X. Yang, S. Li. Learning adaptive and view-invariant vision transformer for real-time UAV tracking. In Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, Article number 1140, 2024."},{"key":"1634_CR6","doi-asserted-by":"publisher","first-page":"7588","DOI":"10.1609\/aaai.v38i7.28591","volume-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence","author":"Y Zheng","year":"2024","unstructured":"Y. Zheng, B. Zhong, Q. Liang, Z. Mo, S. Zhang, X. Li. ODTrack: Online dense temporal token learning for visual tracking. In Proceedings of the 38th AAAI Conference on Artificial Intelligence, Vancouver, Canada, pp. 7588\u20137596, 2024. DOI: https:\/\/doi.org\/10.1609\/aaai.v38i7.28591."},{"issue":"9","key":"1634_CR7","doi-asserted-by":"publisher","first-page":"1834","DOI":"10.1109\/TPAMI.2014.2388226","volume":"37","author":"Y Wu","year":"2015","unstructured":"Y. Wu, J. Lim, M. H. Yang. Object tracking benchmark. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 37, no. 9, pp. 1834\u20131848, 2015. DOI: https:\/\/doi.org\/10.1109\/TPAMI.2014.2388226.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"11","key":"1634_CR8","doi-asserted-by":"publisher","first-page":"10551","DOI":"10.1109\/TCSVT.2024.3409898","volume":"34","author":"J Zhu","year":"2024","unstructured":"J. Zhu, X. Chen, P. Zhang, X. Wang, D. Wang, W. Zhao, H. Lu. SRRT: Exploring search region regulation for visual object tracking. IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 11, pp. 10551\u201310563, 2024. DOI: https:\/\/doi.org\/10.1109\/TCSVT.2024.3409898.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"1634_CR9","doi-asserted-by":"publisher","first-page":"19258","DOI":"10.1109\/CVPR52733.2024.01822","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"W Cai","year":"2024","unstructured":"W. Cai, Q. Liu, Y. Wang. HIPTrack: Visual tracking with historical prompts. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 19258\u201319267, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.01822."},{"key":"1634_CR10","doi-asserted-by":"publisher","first-page":"19113","DOI":"10.1109\/CVPR52733.2024.01808","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"F Xie","year":"2024","unstructured":"F. Xie, Z. Wang, C. Ma. DiffusionTrack: Point set diffusion model for visual object tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 19113\u201319124, 2024. DOI: https:\/\/doi.org\/10.1109\/CVPR52733.2024.01808."},{"key":"1634_CR11","doi-asserted-by":"publisher","first-page":"1420","DOI":"10.1109\/CVPR.2016.158","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition","author":"R Tao","year":"2016","unstructured":"R. Tao, E. Gavves, A. W. M. Smeulders. Siamese instance search for tracking. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, pp. 1420\u20131429, 2016. DOI: https:\/\/doi.org\/10.1109\/CVPR.2016.158."},{"key":"1634_CR12","doi-asserted-by":"publisher","first-page":"850","DOI":"10.1007\/978-3-319-48881-3_56","volume-title":"Proceedings of European Conference on Computer Vision","author":"L Bertinetto","year":"2016","unstructured":"L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, P. H. S. Torr. Fully-convolutional Siamese networks for object tracking. In Proceedings of European Conference on Computer Vision, Amsterdam, The Netherlands, pp. 850\u2013865, 2016. DOI: https:\/\/doi.org\/10.1007\/978-3-319-48881-3_56."},{"key":"1634_CR13","doi-asserted-by":"publisher","first-page":"7944","DOI":"10.1109\/CVPR.2019.00814","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"H Fan","year":"2019","unstructured":"H. Fan, H. Ling. Siamese cascaded region proposal networks for real-time visual tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, pp. 7944\u20137953, 2019. DOI: https:\/\/doi.org\/10.1109\/CVPR.2019.00814."},{"key":"1634_CR14","doi-asserted-by":"publisher","first-page":"1781","DOI":"10.1109\/ICCV.2017.196","volume-title":"Proceedings of IEEE International Conference on Computer Vision","author":"Q Guo","year":"2017","unstructured":"Q. Guo, W. Feng, C. Zhou, R. Huang, L. Wan, S. Wang. Learning dynamic Siamese network for visual object tracking. In Proceedings of IEEE International Conference on Computer Vision, Venice, Italy, pp. 1781\u20131789, 2017. DOI: https:\/\/doi.org\/10.1109\/ICCV.2017.196."},{"key":"1634_CR15","doi-asserted-by":"publisher","first-page":"4277","DOI":"10.1109\/CVPR.2019.00441","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"B Li","year":"2019","unstructured":"B. Li, W. Wu, Q. Wang, F. Zhang, J. Xing, J. Yan. SiamRPN++: Evolution of Siamese visual tracking with very deep networks. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, pp. 4277\u20134286, 2019. DOI: https:\/\/doi.org\/10.1109\/CVPR.2019.00441."},{"key":"1634_CR16","doi-asserted-by":"publisher","first-page":"8971","DOI":"10.1109\/CVPR.2018.00935","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"B Li","year":"2018","unstructured":"B. Li, J. Yan, W. Wu, Z. Zhu, X. Hu. High performance visual tracking with Siamese region proposal network. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, pp. 8971\u20138980, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00935."},{"key":"1634_CR17","doi-asserted-by":"publisher","first-page":"771","DOI":"10.1007\/978-3-03058589-1_46","volume-title":"Proceedings of the 16th European Conference on Computer Vision","author":"Z Zhang","year":"2020","unstructured":"Z. Zhang, H. Peng, J. Fu, B. Li, W. Hu. OCEAN: Object-aware anchor-free tracking. In Proceedings of the 16th European Conference on Computer Vision, Glasgow, UK, pp. 771\u2013787, 2020. DOI: https:\/\/doi.org\/10.1007\/978-3-03058589-1_46."},{"key":"1634_CR18","doi-asserted-by":"publisher","first-page":"6181","DOI":"10.1109\/ICCV.2019.00628","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"G Bhat","year":"2019","unstructured":"G. Bhat, M. Danelljan, L. Van Gool, R. Timofte. Learning discriminative model prediction for tracking. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea, pp. 6181\u20136190, 2019. DOI: https:\/\/doi.org\/10.1109\/ICCV.2019.00628."},{"key":"1634_CR19","doi-asserted-by":"publisher","first-page":"4655","DOI":"10.1109\/CVPR.2019.00479","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"M Danelljan","year":"2019","unstructured":"M. Danelljan, G. Bhat, F. S. Khan, M. Felsberg. ATOM: Accurate tracking by overlap maximization. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, pp. 4655\u20134664, 2019. DOI: https:\/\/doi.org\/10.1109\/CVPR.2019.00479."},{"key":"1634_CR20","volume-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems","author":"T Vu","year":"2019","unstructured":"T. Vu, H. Jang, T. X. Pham, C. D. Yoo. Cascade RPN: Delving into high-quality region proposal network with adaptive convolution. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Vancouver, Canada, Article number 128, 2019."},{"key":"1634_CR21","doi-asserted-by":"publisher","first-page":"2980","DOI":"10.1109\/ICCV.2017.322","volume-title":"Proceedings of IEEE International Conference on Computer Vision","author":"K He","year":"2017","unstructured":"K. He, G. Gkioxari, P. Dollar, R. Girshick. Mask R-CNN. In Proceedings of IEEE International Conference on Computer Vision, Venice, Italy, pp. 2980\u20132988, 2017. DOI: https:\/\/doi.org\/10.1109\/ICCV.2017.322."},{"key":"1634_CR22","doi-asserted-by":"publisher","first-page":"89","DOI":"10.1007\/978-3-030-01225-0_6","volume-title":"Proceedings of the 15th European Conference on Computer Vision","author":"I Jung","year":"2018","unstructured":"I. Jung, J. Son, M. Baek, B. Han. Real-time MDNet. In Proceedings of the 15th European Conference on Computer Vision, Munich, Germany, pp. 89\u2013104, 2018. DOI: https:\/\/doi.org\/10.1007\/978-3-030-01225-0_6."},{"key":"1634_CR23","doi-asserted-by":"publisher","first-page":"3638","DOI":"10.1109\/CVPR.2019.00376","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"G Wang","year":"2019","unstructured":"G. Wang, C. Luo, Z. Xiong, W. Zeng. SPM-tracker: Series-parallel matching for real-time visual object tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, pp. 3638\u20133647, 2019. DOI: https:\/\/doi.org\/10.1109\/CVPR.2019.00376."},{"key":"1634_CR24","doi-asserted-by":"publisher","first-page":"749","DOI":"10.1007\/978-3-319-46448-0_45","volume-title":"Proceedings of the 14th European Conference on Computer Vision","author":"D Held","year":"2016","unstructured":"D. Held, S. Thrun, S. Savarese. Learning to track at 100 FPS with deep regression networks. In Proceedings of the 14th European Conference on Computer Vision, Amsterdam, The Netherlands, pp. 749\u2013765, 2016. DOI: https:\/\/doi.org\/10.1007\/978-3-319-46448-0_45."},{"key":"1634_CR25","doi-asserted-by":"publisher","first-page":"11037","DOI":"10.1609\/aaai.v34i07.6758","volume-title":"Proceedings of the 34th AAAI Conference on Artificial Intelligence","author":"L Huang","year":"2020","unstructured":"L. Huang, X. Zhao, K. Huang. GlobalTrack: A simple and strong baseline for long-term tracking. In Proceedings of the 34th AAAI Conference on Artificial Intelligence, New York, USA, pp. 11037\u201311044, 2020. DOI: https:\/\/doi.org\/10.1609\/aaai.v34i07.6758."},{"key":"1634_CR26","doi-asserted-by":"publisher","first-page":"6568","DOI":"10.1109\/ICCV.2019.00667","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"K Duan","year":"2019","unstructured":"K. Duan, S. Bai, L. Xie, H. Qi, Q. Huang, Q. Tian. CenterNet: Keypoint triplets for object detection. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea, pp. 6568\u20136577, 2019. DOI: https:\/\/doi.org\/10.1109\/ICCV.2019.00667."},{"key":"1634_CR27","doi-asserted-by":"publisher","first-page":"765","DOI":"10.1007\/978-3-030-01264-9_45","volume-title":"Proceedings of the 15th European Conference on Computer Vision","author":"H Law","year":"2018","unstructured":"H. Law, J. Deng. CornerNet: Detecting objects as paired keypoints. In Proceedings of the 15th European Conference on Computer Vision, Munich, Germany, pp. 765\u2013781, 2018. DOI: https:\/\/doi.org\/10.1007\/978-3-030-01264-9_45."},{"key":"1634_CR28","doi-asserted-by":"publisher","first-page":"9626","DOI":"10.1109\/ICCV.2019.00972","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"Z Tian","year":"2019","unstructured":"Z. Tian, C. Shen, H. Chen, T. He. FCOS: Fully convolutional one-stage object detection. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea, pp. 9626\u20139635, 2019. DOI: https:\/\/doi.org\/10.1109\/ICCV.2019.00972."},{"key":"1634_CR29","doi-asserted-by":"publisher","first-page":"516","DOI":"10.1145\/2964284.2967274","volume-title":"Proceedings of the 24th ACM International Conference on Multimedia","author":"J Yu","year":"2016","unstructured":"J. Yu, Y. Jiang, Z. Wang, Z. Cao, T. Huang. UnitBox: An advanced object detection network. In Proceedings of the 24th ACM International Conference on Multimedia, Amsterdam, The Netherlands, pp. 516\u2013520, 2016. DOI: https:\/\/doi.org\/10.1145\/2964284.2967274."},{"key":"1634_CR30","doi-asserted-by":"publisher","first-page":"6667","DOI":"10.1109\/CVPR42600.2020.00670","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Z Chen","year":"2020","unstructured":"Z. Chen, B. Zhong, G. Li, S. Zhang, R. Ji. Siamese box adaptive network for visual tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 6667\u20136676, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00670."},{"key":"1634_CR31","doi-asserted-by":"publisher","first-page":"12549","DOI":"10.1609\/aaai.v34i07.6944","volume-title":"Proceedings of the 34th AAAI Conference on Artificial Intelligence","author":"Y Xu","year":"2020","unstructured":"Y. Xu, Z. Wang, Z. Li, Y. Yuan, G. Yu. SiamFC++: Towards robust and accurate visual tracking with target estimation guidelines. In Proceedings of the 34th AAAI Conference on Artificial Intelligence, New York, USA, pp. 12549\u201312556, 2020. DOI: https:\/\/doi.org\/10.1609\/aaai.v34i07.6944."},{"key":"1634_CR32","doi-asserted-by":"publisher","first-page":"6268","DOI":"10.1109\/CVPR42600.2020.00630","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"D Guo","year":"2020","unstructured":"D. Guo, J. Wang, Y. Cui, Z. Wang, S. Chen. SiamCAR: Siamese fully convolutional classification and regression for visual tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 6268\u20136276, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00630."},{"issue":"12","key":"1634_CR33","doi-asserted-by":"publisher","first-page":"11446","DOI":"10.1109\/TPAMI.2024.3457886","volume":"46","author":"W Hu","year":"2024","unstructured":"W. Hu, S. Wang, Z. Zhou, J. Gao, Y. Li, S. Maybank. One-stage anchor-free online multiple target tracking with deformable local attention and task-aware prediction. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 11446\u201311463, 2024. DOI: https:\/\/doi.org\/10.1109\/TPAMI.2024.3457886.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1634_CR34","doi-asserted-by":"publisher","first-page":"2999","DOI":"10.1109\/ICCV.2017.324","volume-title":"Proceedings of IEEE International Conference on Computer Vision","author":"T Y Lin","year":"2017","unstructured":"T. Y. Lin, P. Goyal, R. Girshick, K. He, P. Doll\u00e1r. Focal loss for dense object detection. In Proceedings of IEEE International Conference on Computer Vision, Venice, Italy, pp. 2999\u20133007, 2017. DOI: https:\/\/doi.org\/10.1109\/ICCV.2017.324."},{"key":"1634_CR35","doi-asserted-by":"publisher","first-page":"7151","DOI":"10.1109\/CVPR.2018.00747","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"H Zhang","year":"2018","unstructured":"H. Zhang, K. Dana, J. Shi, Z. Zhang, X. Wang, A. Tyagi, A. Agrawal. Context encoding for semantic segmentation. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, pp. 7151\u20137160, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00747."},{"key":"1634_CR36","doi-asserted-by":"publisher","first-page":"13319","DOI":"10.1109\/ICCV48922.2021.01309","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"Z Zhang","year":"2021","unstructured":"Z. Zhang, Y. Liu, X. Wang, B. Li, W. Hu. Learn to match: Automatic matching network design for visual tracking. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Montreal, Canada, pp. 13319\u201313328, 2021. DOI: https:\/\/doi.org\/10.1109\/ICCV48922.2021.01309."},{"key":"1634_CR37","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2020.106079","volume-title":"Knowledge-Based Systems","author":"K Yang","year":"2020","unstructured":"K. Yang, Z. He, Z. Zhou, N. Fan. SiamATT: Siamese attention network for visual tracking. Knowledge-Based Systems, vol. 203, Article number 106079, 2020. DOI: https:\/\/doi.org\/10.1016\/j.knosys.2020.106079."},{"issue":"3","key":"1634_CR38","doi-asserted-by":"publisher","first-page":"3072","DOI":"10.1109\/TPAMI.2022.3172932","volume":"45","author":"W Hu","year":"2023","unstructured":"W. Hu, Q. Wang, L. Zhang, L. Bertinetto, P. H. S. Torr. SiamMask: A framework for fast online object tracking and segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 3, pp. 3072\u20133089, 2023. DOI: https:\/\/doi.org\/10.1109\/TPAMI.2022.3172932.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1634_CR39","doi-asserted-by":"publisher","first-page":"1199","DOI":"10.1109\/CVPR.2018.00131","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"F Sung","year":"2018","unstructured":"F. Sung, Y. Yang, L. Zhang, T. Xiang, P. H. S. Torr, T. M. Hospedales. Learning to compare: Relation network for few-shot learning. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, pp. 1199\u20131208, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00131."},{"key":"1634_CR40","doi-asserted-by":"publisher","first-page":"6947","DOI":"10.1109\/CVPR42600.2020.00698","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Y Zhang","year":"2020","unstructured":"Y. Zhang, Z. Wu, H. Peng, S. Lin. A transductive approach for video object segmentation. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 6947\u20136956, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00698."},{"key":"1634_CR41","doi-asserted-by":"publisher","first-page":"3942","DOI":"10.1609\/aaai.v32i1.11671","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence","author":"E Perez","year":"2018","unstructured":"E. Perez, F. Strub, H. de Vries, V. Dumoulin, A. Courville. FiLM: Visual reasoning with a general conditioning layer. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence, New Orleans, USA, pp. 3942\u20133951, 2018. DOI: https:\/\/doi.org\/10.1609\/aaai.v32i1.11671."},{"issue":"6","key":"1634_CR42","doi-asserted-by":"publisher","first-page":"7478","DOI":"10.1109\/TNNLS.2022.3227717","volume":"35","author":"Y Liu","year":"2024","unstructured":"Y. Liu, Y. Zhang, Y. Wang, F. Hou, J. Yuan, J. Tian, Y. Zhang, Z. Shi, J. Fan, Z. He. A survey of visual transformers. IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 6, pp. 7478\u20137498, 2024. DOI: https:\/\/doi.org\/10.1109\/TNNLS.2022.3227717.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"1634_CR43","volume-title":"Proceedings of the 9th International Conference on Learning Representations","author":"A Dosovitskiy","year":"2021","unstructured":"A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, N. Houlsby. An image is worth 16\u00d716 words: Transformers for image recognition at scale. In Proceedings of the 9th International Conference on Learning Representations, 2021."},{"key":"1634_CR44","doi-asserted-by":"publisher","first-page":"2317","DOI":"10.1109\/CVPR42600.2020.00239","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"T Verelst","year":"2020","unstructured":"T. Verelst, T. Tuytelaars. Dynamic convolutions: Exploiting spatial sparsity for faster inference. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 2317\u20132326, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00239."},{"key":"1634_CR45","doi-asserted-by":"publisher","first-page":"465","DOI":"10.1007\/978-3-030-58555-6_28","volume-title":"Proceedings of the 16th European Conference on Computer Vision","author":"X Chu","year":"2020","unstructured":"X. Chu, T. Zhou, B. Zhang, J. Li. Fair DARTS: Eliminating unfair advantages in differentiable architecture search. In Proceedings of the 16th European Conference on Computer Vision, Glasgow, UK, pp. 465\u2013480, 2020. DOI: https:\/\/doi.org\/10.1007\/978-3-030-58555-6_28."},{"key":"1634_CR46","volume-title":"Proceedings of the 7th International Conference on Learning Representation","author":"H Liu","year":"2019","unstructured":"H. Liu, K. Simonyan, Y. Yang. DARTS: Differentiable architecture search. In Proceedings of the 7th International Conference on Learning Representation, New Orleans, USA, 2019."},{"key":"1634_CR47","volume-title":"Proceedings of the 8th International Conference on Learning Representations","author":"B E Bejnordi","year":"2020","unstructured":"B. E. Bejnordi, T. Blankevoort, M. Welling. Batch-shaping for learning conditional channel gated networks. In Proceedings of the 8th International Conference on Learning Representations, Addis Ababa, Ethiopia, 2020."},{"key":"1634_CR48","doi-asserted-by":"publisher","first-page":"7464","DOI":"10.1109\/CVPR.2017.789","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition","author":"E Real","year":"2017","unstructured":"E. Real, J. Shlens, S. Mazzocchi, X. Pan, V. Vanhoucke. YouTube-BoundingBoxes: A large high-precision human-annotated data set for object detection in video. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, pp. 7464\u20137473, 2017. DOI: https:\/\/doi.org\/10.1109\/CVPR.2017.789."},{"issue":"3","key":"1634_CR49","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, Li F. F. ImageNet large scale visual recognition challenge. International Journal of Computer Vision, vol. 115, no. 3, pp. 211\u2013252, 2015. DOI: https:\/\/doi.org\/10.1007\/s11263-015-0816-y.","journal-title":"International Journal of Computer Vision"},{"key":"1634_CR50","doi-asserted-by":"publisher","first-page":"740","DOI":"10.1007\/978-3-319-10602-1_48","volume-title":"Proceedings of the 13th European Conference on Computer Vision","author":"T Y Lin","year":"2014","unstructured":"T. Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll\u00e1r, C. L. Zitnick. Microsoft COCO: Common objects in context. In Proceedings of the 13th European Conference on Computer Vision, Z\u00fcrich, Switzerland, pp. 740\u2013755, 2014. DOI: https:\/\/doi.org\/10.1007\/978-3-319-10602-1_48."},{"issue":"5","key":"1634_CR51","doi-asserted-by":"publisher","first-page":"1562","DOI":"10.1109\/TPAMI.2019.2957464","volume":"43","author":"L Huang","year":"2021","unstructured":"L. Huang, X. Zhao, K. Huang. GOT-10K: A large high-diversity benchmark for generic object tracking in the wild. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 5, pp. 1562\u20131577, 2021. DOI: https:\/\/doi.org\/10.1109\/TPAMI.2019.2957464.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1634_CR52","doi-asserted-by":"publisher","first-page":"5369","DOI":"10.1109\/CVPR.2019.00552","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"H Fan","year":"2019","unstructured":"H. Fan, L. Lin, F. Yang, P. Chu, G. Deng, S. Yu, H. Bai, Y. Xu, C. Liao, H. Ling. LaSOT: A high-quality benchmark for large-scale single object tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, pp. 5369\u20135378, 2019. DOI: https:\/\/doi.org\/10.1109\/CVPR.2019.00552."},{"key":"1634_CR53","doi-asserted-by":"publisher","first-page":"6161","DOI":"10.1109\/ICCV.2019.00626","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"P Li","year":"2019","unstructured":"P. Li, B. Chen, W. Ouyang, D. Wang, X. Yang, H. Lu. GradNet: Gradient-guided network for visual object tracking. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea, pp. 6161\u20136170, 2019. DOI: https:\/\/doi.org\/10.1109\/ICCV.2019.00626."},{"key":"1634_CR54","doi-asserted-by":"publisher","first-page":"153","DOI":"10.1007\/978-3-030-01240-3_10","volume-title":"Proceedings of the 15th European Conference on Computer Vision","author":"T Yang","year":"2018","unstructured":"T. Yang, A. B. Chan. Learning dynamic memory networks for object tracking. In Proceedings of the 15th European Conference on Computer Vision, Munich, Germany, pp. 153\u2013169, 2018. DOI: https:\/\/doi.org\/10.1007\/978-3-030-01240-3_10."},{"key":"1634_CR55","doi-asserted-by":"publisher","first-page":"6577","DOI":"10.1109\/CVPR42600.2020.00661","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"P Voigtlaender","year":"2020","unstructured":"P. Voigtlaender, J. Luiten, P. H. S. Torr, B. Leibe. Siam R-CNN: Visual tracking by re-detection. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 6577\u20136587, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00661."},{"key":"1634_CR56","doi-asserted-by":"publisher","first-page":"6727","DOI":"10.1109\/CVPR42600.2020.00676","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Y Yu","year":"2020","unstructured":"Y. Yu, Y. Xiong, W. Huang, M. R. Scott. Deformable Siamese attention networks for visual object tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 6727\u20136736, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00676."},{"issue":"4","key":"1634_CR57","doi-asserted-by":"publisher","first-page":"5158","DOI":"10.1109\/tpami.2022.3195759","volume":"45","author":"Z Chen","year":"2023","unstructured":"Z. Chen, B. Zhong, G. Li, S. Zhang, R. Ji, Z. Tang, X. Li. SiamBAN: Target-aware tracking with Siamese box adaptive network. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 4, pp. 5158\u20135173, 2023. DOI: https:\/\/doi.org\/10.1109\/tpami.2022.3195759.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1634_CR58","doi-asserted-by":"publisher","first-page":"9538","DOI":"10.1109\/CVPR46437.2021.00942","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"D Guo","year":"2021","unstructured":"D. Guo, Y. Shao, Y. Cui, Z. Wang, L. Zhang, C. Shen. Graph attention tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 9538\u20139547, 2021. DOI: https:\/\/doi.org\/10.1109\/CVPR46437.2021.00942."},{"key":"1634_CR59","doi-asserted-by":"publisher","first-page":"15175","DOI":"10.1109\/CVPR46437.2021.01493","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"B Yan","year":"2021","unstructured":"B. Yan, H. Peng, K. Wu, D. Wang, J. Fu, H. Lu. Light-Track: Finding lightweight neural networks for object tracking via one-shot architecture search. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 15175\u201315184, 2021. DOI: https:\/\/doi.org\/10.1109\/CVPR46437.2021.01493."},{"key":"1634_CR60","doi-asserted-by":"publisher","first-page":"9578","DOI":"10.1109\/ICCV51070.2023.00881","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"B Kang","year":"2023","unstructured":"B. Kang, X. Chen, D. Wang, H. Peng, H. Lu. Exploring lightweight hierarchical vision transformers for efficient visual tracking. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Paris, France, pp. 9578\u20139587, 2023. DOI: https:\/\/doi.org\/10.1109\/ICCV51070.2023.00881."},{"key":"1634_CR61","doi-asserted-by":"publisher","first-page":"8122","DOI":"10.1109\/CVPR46437.2021.00803","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"X Chen","year":"2021","unstructured":"X. Chen, B. Yan, J. Zhu, D. Wang, X. Yang, H. Lu. Transformer tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 8122\u20138131, 2021. DOI: https:\/\/doi.org\/10.1109\/CVPR46437.2021.00803."},{"key":"1634_CR62","doi-asserted-by":"publisher","first-page":"10428","DOI":"10.1109\/ICCV48922.2021.01028","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"B Yan","year":"2021","unstructured":"B. Yan, H. Peng, J. Fu, D. Wang, H. Lu. Learning spatio-temporal transformer for visual tracking. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Montreal, Canada, pp. 10428\u201310437, 2021. DOI: https:\/\/doi.org\/10.1109\/ICCV48922.2021.01028."},{"key":"1634_CR63","doi-asserted-by":"publisher","first-page":"461","DOI":"10.1007\/978-3-031-25085-9_26","volume-title":"Proceedings of European Conference on Computer Vision","author":"X Chen","year":"2023","unstructured":"X. Chen, B. Kang, D. Wang, D. Li, H. Lu. Efficient visual tracking via hierarchical cross-attention transformer. In Proceedings of European Conference on Computer Vision, Tel Aviv, Israel, pp. 461\u2013477, 2023. DOI: https:\/\/doi.org\/10.1007\/978-3-031-25085-9_26."},{"issue":"4","key":"1634_CR64","doi-asserted-by":"publisher","first-page":"6478","DOI":"10.1109\/TNNLS.2024.3402994","volume":"36","author":"H Liu","year":"2025","unstructured":"H. Liu, Y. Cai, B. Fan, J. Xu. ScalableTrack: Scalable one-stream tracking via alternating learning. IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 4, pp. 6478\u20136491, 2025. DOI: https:\/\/doi.org\/10.1109\/tnnls.2024.3402994.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"1634_CR65","doi-asserted-by":"publisher","first-page":"5000","DOI":"10.1109\/CVPR.2017.531","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition","author":"J Valmadre","year":"2017","unstructured":"J. Valmadre, L. Bertinetto, J. Henriques, A. Vedaldi, P. H. S. Torr. End-to-end representation learning for correlation filter based tracking. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, pp. 5000\u20135008, 2017. DOI: https:\/\/doi.org\/10.1109\/CVPR.2017.531."},{"key":"1634_CR66","doi-asserted-by":"publisher","first-page":"6631","DOI":"10.1109\/CVPR.2017.733","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition","author":"M Danelljan","year":"2017","unstructured":"M. Danelljan, G. Bhat, F. S. Khan, M. Felsberg. ECO: Efficient convolution operators for tracking. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, USA, pp. 6631\u20136639, 2017. DOI: https:\/\/doi.org\/10.1109\/CVPR.2017.733."},{"key":"1634_CR67","doi-asserted-by":"publisher","first-page":"4904","DOI":"10.1109\/CVPR.2018.00515","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"F Li","year":"2018","unstructured":"F. Li, C. Tian, W. Zuo, L. Zhang, M. H. Yang. Learning spatial-temporal regularized correlation filters for visual tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, pp. 4904\u20134913, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00515."},{"issue":"11","key":"1634_CR68","doi-asserted-by":"publisher","first-page":"5596","DOI":"10.1109\/TIP.2019.2919201","volume":"28","author":"T Xu","year":"2019","unstructured":"T. Xu, Z. H. Feng, X. J. Wu, J. Kittler. Learning adaptive discriminative correlation filters via temporal consistency preserving spatial feature selection for robust visual object tracking. IEEE Transactions on Image Processing, vol. 28, no. 11, pp. 5596\u20135609, 2019. DOI: https:\/\/doi.org\/10.1109\/TIP.2019.2919201.","journal-title":"IEEE Transactions on Image Processing"},{"key":"1634_CR69","doi-asserted-by":"publisher","first-page":"1308","DOI":"10.1109\/CVPR.2019.00140","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"N Wang","year":"2019","unstructured":"N. Wang, Y. Song, C. Ma, W. Zhou, W. Liu, H. Li. Unsupervised deep tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, pp. 1308\u20131317, 2019. DOI: https:\/\/doi.org\/10.1109\/CVPR.2019.00140."},{"key":"1634_CR70","volume-title":"Proceedings of the 30th British Machine Vision Conference","author":"A S Tripathi","year":"2019","unstructured":"A. S. Tripathi, M. Danelljan, L. Van Gool, R. Timofte. Tracking the known and the unknown by leveraging semantic information. In Proceedings of the 30th British Machine Vision Conference, Cardiff, UK, Article number 292, 2019."},{"key":"1634_CR71","doi-asserted-by":"publisher","first-page":"3194","DOI":"10.1109\/TMM.2023.3307939","volume":"26","author":"K Nai","year":"2024","unstructured":"K. Nai, S. Chen. Learning a novel ensemble tracker for robust visual tracking. IEEE Transactions on Multimedia, vol. 26, pp. 3194\u20133206, 2024. DOI: https:\/\/doi.org\/10.1109\/TMM.2023.3307939.","journal-title":"IEEE Transactions on Multimedia"},{"key":"1634_CR72","doi-asserted-by":"publisher","first-page":"7131","DOI":"10.1109\/CVPR42600.2020.00716","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"A Lukezic","year":"2020","unstructured":"A. Lukezic, J. Matas, M. Kristan. D3S \u2013 a discriminative single shot segmentation tracker. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 7131\u20137140, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00716."},{"key":"1634_CR73","doi-asserted-by":"publisher","first-page":"7181","DOI":"10.1109\/CVPR42600.2020.00721","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"M Danelljan","year":"2020","unstructured":"M. Danelljan, L. Van Gool, R. Timofte. Probabilistic regression for visual tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 7181\u20137190, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00721."},{"key":"1634_CR74","doi-asserted-by":"publisher","first-page":"205","DOI":"10.1007\/978-3-030-58592-1_13","volume-title":"Proceedings of the 16th European Conference on Computer Vision","author":"G Bhat","year":"2020","unstructured":"G. Bhat, M. Danelljan, L. Van Gool, R. Timofte. Know your surroundings: Exploiting scene information for object tracking. In Proceedings of the 16th European Conference on Computer Vision, Glasgow, UK, pp. 205\u2013221, 2020. DOI: https:\/\/doi.org\/10.1007\/978-3-030-58592-1_13."},{"key":"1634_CR75","doi-asserted-by":"publisher","first-page":"4293","DOI":"10.1109\/CVPR.2016.465","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition","author":"H Nam","year":"2016","unstructured":"H. Nam, B. Han. Learning multi-domain convolutional neural networks for visual tracking. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, pp. 4293\u20134302, 2016. DOI: https:\/\/doi.org\/10.1109\/CVPR.2016.465."},{"key":"1634_CR76","doi-asserted-by":"publisher","first-page":"8990","DOI":"10.1109\/CVPR.2018.00937","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Y Song","year":"2018","unstructured":"Y. Song, C. Ma, X. Wu, L. Gong, L. Bao, W. Zuo, C. Shen, R. W. H. Lau, M. H. Yang. VITAL: Visual tracking via adversarial learning. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, USA, pp. 8990\u20138999, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00937."},{"key":"1634_CR77","doi-asserted-by":"publisher","first-page":"4644","DOI":"10.1109\/CVPR.2019.00478","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"J Gao","year":"2019","unstructured":"J. Gao, T. Zhang, C. Xu. Graph convolutional tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, pp. 4644\u20134654, 2019. DOI: https:\/\/doi.org\/10.1109\/CVPR.2019.00478."},{"key":"1634_CR78","doi-asserted-by":"publisher","first-page":"587","DOI":"10.1007\/978-3-030-01219-9_35","volume-title":"Proceedings of the 15th European Conference on Computer Vision","author":"E Park","year":"2018","unstructured":"E. Park, A. C. Berg. Meta-tracker: Fast and robust online adaptation for visual object trackers. In Proceedings of the 15th European Conference on Computer Vision, Munich, Germany, pp. 587\u2013604, 2018. DOI: https:\/\/doi.org\/10.1007\/978-3-030-01219-9_35."},{"key":"1634_CR79","doi-asserted-by":"publisher","first-page":"6717","DOI":"10.1109\/CVPR42600.2020.00675","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"T Yang","year":"2020","unstructured":"T. Yang, P. Xu, R. Hu, H. Chai, A. B. Chan. ROAM: Recurrently optimizing tracking model. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 6717\u20136726, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00675."},{"key":"1634_CR80","doi-asserted-by":"publisher","first-page":"6288","DOI":"10.1109\/CVPR42600.2020.00632","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"G Wang","year":"2020","unstructured":"G. Wang, C. Luo, X. Sun, Z. Xiong, W. Zeng. Tracking by instance detection: A meta-learning approach. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, pp. 6288\u20136297, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00632."},{"key":"1634_CR81","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1007\/978-3-030-11009-3_1","volume-title":"Proceedings of European Conference on Computer Vision Workshops","author":"M Kristan","year":"2019","unstructured":"M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, L. \u010c. Zajc, T. Voj\u00edr, G. Bhat, A. Luke\u017ei\u010d, A. Eldesokey, G. Fern\u00e1ndez, \u00c1. Garc\u00eda-Mart\u00edn, \u00c1. Iglesias-Arias, A. A. Alatan, A. Gonz\u00e1lez-Garc\u00eda, A. Petrosino, A. Memarmoghadam, A. Vedaldi, A. Muhi\u010d, A. He, A. Smeulders, A. G. Perera, B. Li, B. Chen, C. Kim, C. Xu, C. Xiong, C. Tian, C. Luo, C. Sun, C. Hao, D. Kim, D. Mishra, D. Chen, D. Wang, D. Wee, E. Gavves, E. Gundogdu, E. Velasco-Salido, F. S. Khan, F. Yang, F. Zhao, F. Li, F. Battistone, G. De Ath, G. R. K. S. Subrahmanyam, G. Bastos, H. Ling, H. K. Galoogahi, H. Lee, H. Li, H. Zhao, H. Fan, H. Zhang, H. Possegger, H. Li, H. Lu, H. Zhi, H. Li, H. Lee, H. J. Chang, I. Drummond, J. Valmadre, J. S. Martin, J. Chahl, J. Y. Choi, J. Li, J. Wang, J. Qi, J. Sung, J. Johnander, J. Henriques, J. Choi, J. van de Weijer, J. R. Herranz, J. M. Mart\u00ednez, J. Kittler, J. Zhuang, J. Gao, K. Grm, L. Zhang, L. Wang, L. Yang, L. Rout, L. Si, L. Bertinetto, L. Chu, M. Che, M. E. Maresca, M. Danelljan, M. H. Yang, M. Abdelpakey, M. Shehata, M. Kang, N. Lee, N. Wang, O. Miksik, P. Moallem, P. Vicente-Mo\u00f1ivar, P. Senna, P. Li, P. Torr, P. M. Raju, R. Qian, Q. Wang, Q. Zhou, Q. Guo, R. Mart\u00edn-Nieto, R. K. Gorthi, R. Tao, R. Bowden, R. Everson, R. Wang, S. Yun, S. Choi, S. Vivas, S. Bai, S. Huang, S. Wu, S. Hadfield, S. Wang, S. Golodetz, T. Ming, T. Xu, T. Zhang, T. Fischer, V. Santopietro, V. \u0160truc, W. Wei, W. Zuo, W. Feng, W. Wu, W. Zou, W. Hu, W. Zhou, W. Zeng, X. Zhang, X. Wu, X. J. Wu, X. Tian, Y. Li, Y. Lu, Y. W. Law, Y. Wu, Y. Demiris, Y. Yang, Y. Jiao, Y. Li, Y. Zhang, Y. Sun, Z. Zhang, Z. Zhu, Z. H. Feng, Z. Wang, Z. He. The sixth visual object tracking VOT2018 challenge results. In Proceedings of European Conference on Computer Vision Workshops, Munich, Germany, pp. 3\u201353, 2019. DOI: https:\/\/doi.org\/10.1007\/978-3-030-11009-3_1."},{"key":"1634_CR82","doi-asserted-by":"publisher","first-page":"2206","DOI":"10.1109\/ICCVW.2019.00276","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision Workshop","author":"M Kristan","year":"2019","unstructured":"M. Kristan, J. Matas, A. Leonardis, M. Felsberg, R. Pflugfelder, J. K. K\u00e4m\u00e4r\u00e4inen, L. C. Zajc, O. Drbohlav, A. Lukezic, A. Berg, A. Eldesokey, J. K\u00e4pyl\u00e4, G. Fern\u00e1ndez, A. Gonzalez-Garcia, A. Memarmoghadam, A. Lu, A. He, A. Varfolomieiev, A. Chan, A. S. Tripathi, A. Smeulders, B. S. Pedasingu, B. X. Chen, B. Zhang, B. Wu, B. Li, B. He, B. Yan, B. Bai, B. Li, B. Li, B. H. Kim, C. Ma, C. Fang, C. Qian, C. Chen, C. Li, C. Zhang, C. Y. Tsai, C. Luo, C. Micheloni, C. Zhang, D. Tao, D. Gupta, D. Song, D. Wang, E. Gavves, E. Yi, F. S. Khan, F. Zhang, F. Wang, F. Zhao, G. De Ath, G. Bhat, G. Chen, G. Wang, G. Li, H. Cevikalp, H. Du, H. Zhao, H. Saribas, H. M. Jung, H. Bai, H. Yu, H. Peng, H. Lu, H. Li, J. Li, J. Li, J. Fu, J. Chen, J. Gao, J. Zhao, J. Tang, J. Li, J. Wu, J. Liu, J. Wang, J. Qi, J. Zhang, J. K. Tsotsos, J. H. Lee, J. van de Weijer, J. Kittler, J. H. Lee, J. Zhuang, K. Zhang, K. Wang, K. Dai, L. Chen, L. Liu, L. Guo, L. Zhang, L. Wang, L. Wang, L. Zhang, L. Wang, L. Zhou, L. Zheng, L. Rout, L. Van Gool, L. Bertinetto, M. Danelljan, M. Dunnhofer, M. Ni, M. Y. Kim, M. Tang, M. H. Yang, N. Paluru, N. Martinel, P. Xu, P. Zhang, P. Zheng, P. Zhang, P. H. S. Torr, Q. Zhang, Q. Wang, Q. Guo, R. Timofte, R. K. Gorthi, R. Everson, R. Han, R. Zhang, S. You, S. C. Zhao, S. Zhao, S. Li, S. Li, S. Ge, S. Bai, S. Guan, T. Xing, T. Xu, T. Yang, T. Zhang, T. Vojir, W. Feng, W. Hu, W. Wang, W. Tang, W. Zeng, W. Liu, X. Chen, X. Qiu, X. Bai, X. J. Wu, X. Yang, X. Chen, X. Li, X. Sun, X. Chen, X. Tian, X. Tang, X. F. Zhu, Y. Huang, Y. Chen, Y. Lian, Y. Gu, Y. Liu, Y. Chen, Y. Zhang, Y. Xu, Y. Wang, Y. Li, Y. Zhou, Y. Dong, Y. Xu, Y. Zhang, Y. Li, Z. Wang Z. Luo, Z. Zhang, Z. H. Feng, Z. He, Z. Song, Z. Chen, Z. Zhang, Z. Wu, Z. Xiong, Z. Huang, Z. Teng, Z. Ni. The seventh visual object tracking VOT2019 challenge results. In Proceedings of IEEE\/CVF International Conference on Computer Vision Workshop, Seoul, Republic of Korea, pp. 2206\u20132241, 2019. DOI: https:\/\/doi.org\/10.1109\/ICCVW.2019.00276."},{"key":"1634_CR83","doi-asserted-by":"publisher","first-page":"310","DOI":"10.1007\/978-3-030-01246-5_19","volume-title":"Proceedings of the 15th European Conference on Computer Vision","author":"M M\u00fcller","year":"2018","unstructured":"M. M\u00fcller, A. Bibi, S. Giancola, S. Alsubaihi, B. Ghanem. TrackingNet: A large-scale dataset and benchmark for object tracking in the wild. In Proceedings of the 15th European Conference on Computer Vision, Munich, Germany, pp. 310\u2013327, 2018. DOI: https:\/\/doi.org\/10.1007\/978-3-030-01246-5_19."},{"key":"1634_CR84","doi-asserted-by":"publisher","first-page":"13758","DOI":"10.1109\/CVPR46437.2021.01355","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"X Wang","year":"2021","unstructured":"X. Wang, X. Shu, Z. Zhang, B. Jiang, Y. Wang, Y. Tian, F. Wu. Towards more flexible and accurate object tracking with natural language: Algorithms and benchmark. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 13758\u201313768, 2021. DOI: https:\/\/doi.org\/10.1109\/CVPR46437.2021.01355."},{"key":"1634_CR85","doi-asserted-by":"publisher","first-page":"2698","DOI":"10.1109\/ICCVW54120.2021.00304","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"M Dunnhofer","year":"2021","unstructured":"M. Dunnhofer, A. Furnari, G. M. Farinella, C. Micheloni. Is first person vision challenging for object tracking? In Proceedings of IEEE\/CVF International Conference on Computer Vision, Montreal, Canada, pp. 2698\u20132710, 2021. DOI: https:\/\/doi.org\/10.1109\/ICCVW54120.2021.00304."},{"key":"1634_CR86","doi-asserted-by":"publisher","first-page":"10714","DOI":"10.1109\/ICCV48922.2021.01056","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"H Fan","year":"2021","unstructured":"H. Fan, H. A. Miththanthaya, H. Harshit, S. R. Rajan, X. Liu, Z. Zou, Y. Lin, H. Ling. Transparent object tracking benchmark. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Montreal, Canada, pp. 10714\u201310723, 2021. DOI: https:\/\/doi.org\/10.1109\/ICCV48922.2021.01056."},{"key":"1634_CR87","doi-asserted-by":"publisher","first-page":"547","DOI":"10.1007\/978-3-030-68238-5_39","volume-title":"Proceedings of International Conference on Computer Vision","author":"M Kristan","year":"2020","unstructured":"M. Kristan, A. Leonardis, J. Matas, M. Felsberg, R. Pflugfelder, J. K. K\u00e4m\u00e4r\u00e4inen, M. Danelljan, L. \u010c. Zajc, A. Luke\u017ei\u010d, O. Drbohlav, L. He, Y. Zhang, S. Yan, J. Yang, G. Fern\u00e1ndez, A. Hauptmann, A. Memarmoghadam, \u00c1. Garc\u00eda-Mart\u00edn, A. Robinson, A. Varfolomieiev, A. H. Gebrehiwot, B. Uzun, B. Yan, B. Li, C. Qian, C. Y. Tsai, C. Micheloni, D. Wang, F. Wang, F. Xie, F. J. Lawin, F. Gustafsson, G. L. Foresti, G. Bhat, G. Chen, H. Ling, H. Zhang, H. Cevikalp, H. Zhao, H. Bai, H. C. Kuchibhotla, H. Saribas, H. Fan, H. Ghanei-Yakhdan, H. Li, H. Peng, H. Lu, H. Li, J. Khaghani, J. Bescos, J. Li, J. Fu, J. Yu, J. Xu, J. Kittler, J. Yin, J. Lee, K. Yu, K. Liu, K. Yang, K. Dai, L. Cheng, L. Zhang, L. Wang, L. Wang, L. Van Gool, L. Bertinetto, M. Dunnhofer, M. Cheng, M. M. Dasari, N. Wang, N. Wang, P. Zhang, P. H. S. Torr, Q. Wang, R. Timofte, R. Krishna S. Gorthi, S. Choi, S. M. Marvasti-Zadeh, S. Zhao, S. Kasaei, S. Qiu, S. Chen, T. B. Sch\u00f6n, T. Xu, W. Lu, W. Hu, W. Zhou, X. Qiu, X. Ke, X. J. Wu, X. Zhang, X. Yang, X. Zhu, Y. Jiang, Y. Wang, Y. Chen, Y. Ye, Y. Li, Y. Yao, Y. Lee, Y. Gu, Z. Wang, Z. Tang, Z. H. Feng, Z. Mai, Z. Zhang, Z. Wu, Z. Ma. The eighth visual object tracking VOT2020 challenge results. In Proceedings of International Conference on Computer Vision, Glasgow, UK, pp. 547\u2013601, 2020. DOI: https:\/\/doi.org\/10.1007\/978-3-030-68238-5_39."},{"key":"1634_CR88","doi-asserted-by":"publisher","first-page":"1134","DOI":"10.1109\/ICCV.2017.128","volume-title":"Proceedings of IEEE international Conference on Computer Vision","author":"H K Galoogahi","year":"2017","unstructured":"H. K. Galoogahi, A. Fagg, C. Huang, D. Ramanan, S. Lucey. Need for speed: A benchmark for higher frame rate object tracking. In Proceedings of IEEE international Conference on Computer Vision, Venice, Italy, pp. 1134\u20131143, 2017. DOI: https:\/\/doi.org\/10.1109\/ICCV.2017.128."},{"issue":"12","key":"1634_CR89","doi-asserted-by":"publisher","first-page":"5630","DOI":"10.1109\/TIP.2015.2482905","volume":"24","author":"P Liang","year":"2015","unstructured":"P. Liang, E. Blasch, H. Ling. Encoding color information for visual tracking: Algorithms and benchmark. IEEE Transactions on Image Processing, vol. 24, no. 12, pp. 5630\u20135644, 2015. DOI: https:\/\/doi.org\/10.1109\/TIP.2015.2482905.","journal-title":"IEEE Transactions on Image Processing"},{"key":"1634_CR90","doi-asserted-by":"publisher","first-page":"445","DOI":"10.1007\/978-3-319-46448-0_27","volume-title":"Proceedings of the 14th European Conference on Computer Vision","author":"M Mueller","year":"2016","unstructured":"M. Mueller, N. Smith, B. Ghanem. A benchmark and simulator for UAV tracking. In Proceedings of the 14th European Conference on Computer Vision, Amsterdam, The Netherlands, pp. 445\u2013461, 2016. DOI: https:\/\/doi.org\/10.1007\/978-3-319-46448-0_27."},{"key":"1634_CR91","doi-asserted-by":"publisher","first-page":"10635","DOI":"10.1609\/aaai.v39i10.33155","volume-title":"Proceedings of the 39th AAAI Conference on Artificial Intelligence","author":"Y Zheng","year":"2025","unstructured":"Y. Zheng, B. Zhong, Q. Liang, N. Li, S. Song. Decoupled spatio-temporal consistency learning for self-supervised tracking. In Proceedings of the 39th AAAI Conference on Artificial Intelligence, Philadelphia, USA, pp. 10635\u201310643, 2025. DOI: https:\/\/doi.org\/10.1609\/aaai.v39i10.33155."},{"issue":"7","key":"1634_CR92","doi-asserted-by":"publisher","first-page":"9186","DOI":"10.1109\/TNNLS.2022.3231537","volume":"35","author":"X Li","year":"2024","unstructured":"X. Li, W. Pei, Y. Wang, Z. He, H. Lu, M. H. Yang. Self-supervised tracking via target-aware data synthesis. IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 7, pp. 9186\u20139197, 2024. DOI: https:\/\/doi.org\/10.1109\/TNNLS.2022.3231537.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"issue":"2","key":"1634_CR93","doi-asserted-by":"publisher","first-page":"400","DOI":"10.1007\/s11263-020-01357-4","volume":"129","author":"N Wang","year":"2021","unstructured":"N. Wang, W. Zhou, Y. Song, C. Ma, W. Liu, H. Li. Unsupervised deep representation learning for real-time tracking. International Journal of Computer Vision, vol. 129, no. 2, pp. 400\u2013418, 2021. DOI: https:\/\/doi.org\/10.1007\/s11263-020-01357-4.","journal-title":"International Journal of Computer Vision"},{"key":"1634_CR94","doi-asserted-by":"publisher","first-page":"10351","DOI":"10.1109\/IROS45743.2020.9341621","volume-title":"Proceedings of IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"W Yuan","year":"2020","unstructured":"W. Yuan, M. Y. Wang, Q. Chen. Self-supervised object tracking with cycle-consistent Siamese networks. In Proceedings of IEEE\/RSJ International Conference on Intelligent Robots and Systems, Las Vegas, USA, pp. 10351\u201310358, 2020. DOI: https:\/\/doi.org\/10.1109\/IROS45743.2020.9341621."},{"key":"1634_CR95","doi-asserted-by":"publisher","first-page":"13526","DOI":"10.1109\/ICCV48922.2021.01329","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"J Zheng","year":"2021","unstructured":"J. Zheng, C. Ma, H. Peng, X. Yang. Learning to track objects from unlabeled videos. In Proceedings of IEEE\/CVF International Conference on Computer Vision, Montreal, Canada, pp. 13526\u201313535, 2021. DOI: https:\/\/doi.org\/10.1109\/ICCV48922.2021.01329."},{"key":"1634_CR96","doi-asserted-by":"publisher","first-page":"1948","DOI":"10.1145\/3394171.3413611","volume-title":"Proceedings of the 28th ACM International Conference on Multimedia","author":"C H Sio","year":"2020","unstructured":"C. H. Sio, Y. J. Ma, H. H. Shuai, J. C. Chen, W. H. Cheng. S2SiamFC: Self-supervised fully convolutional Siamese network for visual tracking. In Proceedings of the 28th ACM International Conference on Multimedia, Seattle, USA, pp. 1948\u20131957, 2020. DOI: https:\/\/doi.org\/10.1145\/3394171.3413611."},{"key":"1634_CR97","doi-asserted-by":"publisher","first-page":"2992","DOI":"10.1109\/CVPR46437.2021.00301","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Q Wu","year":"2021","unstructured":"Q. Wu, J. Wan, A. B. Chan. Progressive unsupervised learning for visual object tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, USA, pp. 2992\u20133001, 2021. DOI: https:\/\/doi.org\/10.1109\/CVPR46437.2021.00301."}],"container-title":["Machine Intelligence Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11633-026-1634-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11633-026-1634-0","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11633-026-1634-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T15:02:01Z","timestamp":1780412521000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11633-026-1634-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,11]]},"references-count":97,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["1634"],"URL":"https:\/\/doi.org\/10.1007\/s11633-026-1634-0","relation":{},"ISSN":["2731-538X","2731-5398"],"issn-type":[{"value":"2731-538X","type":"print"},{"value":"2731-5398","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,11]]},"assertion":[{"value":"23 July 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 January 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 April 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declared that they have no conflicts of interest to this work.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations of conflict of interest"}}]}}