{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,20]],"date-time":"2026-03-20T14:00:13Z","timestamp":1774015213504,"version":"3.50.1"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"8","license":[{"start":{"date-parts":[[2021,5,13]],"date-time":"2021-05-13T00:00:00Z","timestamp":1620864000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,5,13]],"date-time":"2021-05-13T00:00:00Z","timestamp":1620864000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62072122"],"award-info":[{"award-number":["62072122"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61772144"],"award-info":[{"award-number":["61772144"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Education Dept. of Guangdong Province","award":["2019KSYS0092019KSYS009"],"award-info":[{"award-number":["2019KSYS0092019KSYS009"]}]},{"name":"Foreign Science and Technology Cooperation Plan Project of Guangzhou Science Technology and Innovation Commission","award":["201807010059"],"award-info":[{"award-number":["201807010059"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis Comput"],"published-print":{"date-parts":[[2022,8]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>SiamPRN algorithm performs well in visual tracking, but it is easy to drift under occlusion and fast motion scenes because it uses <jats:inline-formula><jats:alternatives><jats:tex-math>$$\\ell _1$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:msub>\n                    <mml:mi>\u2113<\/mml:mi>\n                    <mml:mn>1<\/mml:mn>\n                  <\/mml:msub>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>-smooth loss function to measure the regression location of bounding box. In this paper, we propose a multivariate intersection over union (MIOU) loss in SiamRPN tracking framework. Firstly, MIOU loss includes three geometric factors in regression: the overlap area ratio, the center distance ratio, and the aspect ratio, which can better reflect the coincidence degree of target box and prediction box. Secondly, we improve the definition of aspect ratio loss to avoid gradient explosion, improve the optimization performance of prediction box. Finally, based on SiamPRN tracker, we compared the tracking performance of <jats:inline-formula><jats:alternatives><jats:tex-math>$$\\ell _1$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:msub>\n                    <mml:mi>\u2113<\/mml:mi>\n                    <mml:mn>1<\/mml:mn>\n                  <\/mml:msub>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>-smooth loss, IOU loss, GIOU loss, DIOU loss, and MIOU loss. Experimental results show that the MIOU loss has better target location regression than other loss functions on the OTB2015 and VOT2016 benchmark, especially for the challenges of occlusion, illumination change and fast motion.<\/jats:p>","DOI":"10.1007\/s00371-021-02150-1","type":"journal-article","created":{"date-parts":[[2021,5,13]],"date-time":"2021-05-13T08:03:00Z","timestamp":1620892980000},"page":"2739-2750","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["A multivariate intersection over union of SiamRPN network for visual tracking"],"prefix":"10.1007","volume":"38","author":[{"given":"Zhihui","family":"Huang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huimin","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7070-7031","authenticated-orcid":false,"given":"Jin","family":"Zhan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huakang","family":"Li","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,5,13]]},"reference":[{"key":"2150_CR1","unstructured":"Wang, N.,Yeung, D. Y.: Learning a deep compact image representation for visual tracking. In: Proceedings of the Neural Information Processing Systems (NIPS), pp. 809-817 (2013)"},{"key":"2150_CR2","doi-asserted-by":"crossref","unstructured":"Zhou, X., Xie, L., Zhang, P., Zhang, Y.: An ensemble of deep neural networks for object tracking. In: Proceedings of 2014 IEEE International Conference on Image Processing (ICIP), pp. 843-847 (2014)","DOI":"10.1109\/ICIP.2014.7025169"},{"key":"2150_CR3","unstructured":"Wang, N., Li, S., Gupta, A., Yeung, D.Y.: Transferring rich feature hierarchies for robust visual tracking. arXiv2015 (2015)"},{"key":"2150_CR4","doi-asserted-by":"crossref","unstructured":"Nam, H., Han, B.: Learning multi-domain convolutional neural networks for visual tracking. arXiv 2016(2016)","DOI":"10.1109\/CVPR.2016.465"},{"key":"2150_CR5","doi-asserted-by":"crossref","unstructured":"Tao, R., Gavves, E., Smeulders, A.W.M.: Siamese Instance Search for Tracking. In: Proceedings of IEEE conference on computer vision and pattern recognition (CVPR), pp. 850\u2013865 (2016)","DOI":"10.1109\/CVPR.2016.158"},{"key":"2150_CR6","volume":"8","author":"S Xuan","year":"2020","unstructured":"Xuan, S., Li, S., Zhao, Z., Kou, L., Zhou, Z., Xia, G.: siamese networks with distractor-reduction method for long-term visual object tracking. Pattern Recognit. 8, (2020)","journal-title":"Pattern Recognit."},{"key":"2150_CR7","doi-asserted-by":"crossref","unstructured":"Li, B., Yan, J., Wu, W., Zhu Z., Hu, X.: High performance visual tracking with siamese region proposal network. In: Proceedings of IEEE conference on computer vision and pattern recognition, pp. 8971\u20138980 (2018)","DOI":"10.1109\/CVPR.2018.00935"},{"key":"2150_CR8","doi-asserted-by":"crossref","unstructured":"Grabner, H., Leistner, C., Bischof, H.: Semi-supervised online boosting for robust tracking. In: Proceedings of European Conference on Computer Vision (ECCV), pp. 234\u2013247 (2008)","DOI":"10.1007\/978-3-540-88682-2_19"},{"key":"2150_CR9","doi-asserted-by":"crossref","unstructured":"Babenko, B., Yang, M.H., Belongie, S. Visual tracking with online multiple instance learning. In: Proceedings of IEEE conference on computer vision and pattern recognition (CVPR), pp. 983\u2013990 (2009)","DOI":"10.1109\/CVPR.2009.5206737"},{"issue":"7","key":"2150_CR10","doi-asserted-by":"publisher","first-page":"1409","DOI":"10.1109\/TPAMI.2011.239","volume":"34","author":"Z Kalal","year":"2012","unstructured":"Kalal, Z., Mikolajczyk, K., Matas, J.: Tracking\u2013learning\u2013detection. IEEE Trans. Pattern Anal. Mach. Intell. 34(7), 1409\u20131422 (2012)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2150_CR11","unstructured":"Mei, X., Ling, H.: Robust visual tracking using l1 minimization. In: Proceedings of IEEE International Conference on Computer Vision (ICCV). pp. 1436\u20131443 (2009)"},{"issue":"1","key":"2150_CR12","doi-asserted-by":"publisher","first-page":"314","DOI":"10.1109\/TIP.2012.2202677","volume":"22","author":"D Wang","year":"2013","unstructured":"Wang, D., Lu, H., Yang, M.H.: Online object tracking with sparse prototypes. IEEE Trans. Image Process (TIP) 22(1), 314\u2013325 (2013)","journal-title":"IEEE Trans. Image Process (TIP)"},{"key":"2150_CR13","doi-asserted-by":"crossref","unstructured":"Zhang, T., Liu, S., Xu, C., Yan S., Ghanem Be., Ahuja N., Yang, M.H.: Structural sparse tracking. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 150\u2013158 (2015)","DOI":"10.1109\/CVPR.2015.7298610"},{"key":"2150_CR14","doi-asserted-by":"publisher","first-page":"68","DOI":"10.1016\/j.neucom.2018.01.076","volume":"287","author":"Z Wang","year":"2018","unstructured":"Wang, Z., Ren, J., Zhang, D., Sun, M., Jiang, J.: A deep-learning based feature hybrid framework for spatiotemporal saliency detection inside videos. Neurocomputing 287, 68\u201383 (2018)","journal-title":"Neurocomputing"},{"issue":"1","key":"2150_CR15","doi-asserted-by":"publisher","first-page":"94","DOI":"10.1007\/s12559-017-9529-6","volume":"10","author":"Y Yan","year":"2017","unstructured":"Yan, Y., Ren, J., Zhao, H., Sun, G., Wang, Z., Zheng, J., Marshall, S., Soraghan, J.: Cognitive fusion of thermal and visible imagery for effective detection and tracking of pedestrians in videos. Cognit. Comput. 10(1), 94\u2013104 (2017)","journal-title":"Cognit. Comput."},{"issue":"6","key":"2150_CR16","doi-asserted-by":"publisher","first-page":"3325","DOI":"10.1109\/TGRS.2014.2374218","volume":"53","author":"J Han","year":"2015","unstructured":"Han, J., Zhang, D., Cheng, G., Lei, G., Ren, J.: Object detection in optical remote sensing images based on weakly supervised learning and high-level feature learning. IEEE Trans. Geosci. Remote Sens. 53(6), 3325\u20133337 (2015)","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"2150_CR17","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.neucom.2015.11.044","volume":"185","author":"J Zabalza","year":"2016","unstructured":"Zabalza, J., Ren, J., Zheng, J., Zhao, H., Qing, C., Yang, Z., Du, P., Marshall, S.: Novel segmented stacked autoencoder for effective dimensionality reduction and feature extraction in hyperspectral imaging. Neurocomputing 185, 1\u201310 (2016)","journal-title":"Neurocomputing"},{"key":"2150_CR18","doi-asserted-by":"publisher","first-page":"189","DOI":"10.1016\/j.inffus.2019.02.005","volume":"51","author":"J Tschannerl","year":"2019","unstructured":"Tschannerl, J., Ren, J., Yuen, P., Sun, G., Zhao, H., Yang, Z., Wang, Z., Marshall, S.: MIMR-DGSA: unsupervised hyperspectral band selection based on information theory and a modified discrete gravitational search algorithm. Inf. Fusion 51, 189\u2013200 (2019)","journal-title":"Inf. Fusion"},{"issue":"12","key":"2150_CR19","doi-asserted-by":"publisher","first-page":"3370","DOI":"10.3390\/s20123370","volume":"20","author":"H Xia","year":"2020","unstructured":"Xia, H., Zhang, Y., Yang, M., Zhao, Y.: Visual tracking via deep feature fusion and correlation filters. Sensors 20(12), 3370 (2020)","journal-title":"Sensors"},{"key":"2150_CR20","doi-asserted-by":"crossref","unstructured":"Zhou, X., Xie, L., Zhang, P., et al.: An ensemble of deep neural networks for object tracking. In: Proceedings of the IEEE International Conference on Image Processing (ICIP), pp. 843\u2013847 (2014)","DOI":"10.1109\/ICIP.2014.7025169"},{"key":"2150_CR21","doi-asserted-by":"crossref","unstructured":"Nam, H., Han, B.: Learning multi-domain convolutional neural networks for visual tracking. In: Proceedings of Computer vision and pattern recognition(CVPR), pp. 4293\u20134302 (2016)","DOI":"10.1109\/CVPR.2016.465"},{"key":"2150_CR22","doi-asserted-by":"crossref","unstructured":"Bertinetto, L., Valmadre J., Henriques, J. F., Vedaldi, A., Torr, Philip, H.S.: Fully-Convolutional Siamese Networks for Object Tracking. Proceedings of European Conference on Computer Vision (ECCV).pp.850-865(2016)","DOI":"10.1007\/978-3-319-48881-3_56"},{"key":"2150_CR23","doi-asserted-by":"crossref","unstructured":"Xu, Y., Wang, Z., Li, Z., Yuan, Y., Yu, G.: SiamFC++: towards robust and accurate visual tracking with target estimation guidelines. In: Proceedings of AAAI, pp. 12549\u201312556 (2020)","DOI":"10.1609\/aaai.v34i07.6944"},{"key":"2150_CR24","doi-asserted-by":"crossref","unstructured":"Gao, J., Zhang, T., Xu, C.: Graph convolutional tracking. In: Proceedings of the IEEE conference on computer vision and pattern recognition(CVPR), pp. 4649\u20134659 (2019)","DOI":"10.1109\/CVPR.2019.00478"},{"key":"2150_CR25","doi-asserted-by":"crossref","unstructured":"Li, B., Wu, W., Wang, Q., et al.: SiamRPN++: Evolution of Siamese visual tracking with very deep networks. In: 2019 IEEE\/CVF conference on computer vision and pattern recognition (CVPR). IEEE (2020)","DOI":"10.1109\/CVPR.2019.00441"},{"key":"2150_CR26","doi-asserted-by":"crossref","unstructured":"Zhu, Z., Wang, Q., Li, B., et al.: Distractor-aware Siamese Networks for Visual Object Tracking. In: ECCV2018. Springer, Cham (2018)","DOI":"10.1007\/978-3-030-01240-3_7"},{"key":"2150_CR27","doi-asserted-by":"crossref","unstructured":"Wang, Q., Zhang, L., Bertinetto, L., et al.: Fast online object tracking and segmentation: A unifying approach[C]\/\/Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2019: 1328-1338","DOI":"10.1109\/CVPR.2019.00142"},{"key":"2150_CR28","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Peng, H.: Deeper and wider Siamese networks for real-time visual tracking. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 4591\u20134600 (2019)","DOI":"10.1109\/CVPR.2019.00472"},{"issue":"6","key":"2150_CR29","doi-asserted-by":"publisher","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","volume":"39","author":"S Ren","year":"2015","unstructured":"Ren, S., He, K., Girshick, R., et al.: Faster R-CNN: towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 39(6), 1137\u20131149 (2015)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2150_CR30","doi-asserted-by":"crossref","unstructured":"Song, Y., Ma, C., Wu, X., Gong L., Bao L.,Zuo W., Shen C., Lau, R.W.H., Yang, M.H.: VITAL: Visual tracking via adversarial learning. In: Proceedings of Conference on Computer Vision and Pattern Recognition, pp. 8990\u20138999 (2018)","DOI":"10.1109\/CVPR.2018.00937"},{"key":"2150_CR31","doi-asserted-by":"crossref","unstructured":"Yu, J., Jiang, Y., Wang, Z., et al.: UnitBox: an advanced object detection network. In: Proceedings of the 24th ACM international conference on Multimedia, pp. 516\u2013520 (2016)","DOI":"10.1145\/2964284.2967274"},{"key":"2150_CR32","doi-asserted-by":"crossref","unstructured":"Rezatofighi, H.,Tsoi, N., Gwak, J.Y., Sadeghian, A., Reid, I., Savarese, S.: Generalized intersection over union: a metric and a loss for bounding box regression. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 658\u2013666 (2019)","DOI":"10.1109\/CVPR.2019.00075"},{"key":"2150_CR33","doi-asserted-by":"crossref","unstructured":"Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., Ren, D.: Distance-IoU Loss: faster and better learning for bounding box regression. In: Proceedings of AAAI, pp. 12993\u201313000 (2020)","DOI":"10.1609\/aaai.v34i07.6999"},{"key":"2150_CR34","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., et al.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770\u2013778 (2016)","DOI":"10.1109\/CVPR.2016.90"},{"issue":"2","key":"2150_CR35","doi-asserted-by":"publisher","first-page":"127","DOI":"10.1023\/A:1010091220143","volume":"1","author":"R Rubinstein","year":"1999","unstructured":"Rubinstein, R.: The cross-entropy method for combinatorial and continuous optimization. Methodol. Comput. Appl. Probab. 1(2), 127\u2013190 (1999)","journal-title":"Methodol. Comput. Appl. Probab."},{"key":"2150_CR36","doi-asserted-by":"publisher","first-page":"1834","DOI":"10.1109\/TPAMI.2014.2388226","volume":"37","author":"Y Wu","year":"2015","unstructured":"Wu, Y., Lim, J., Yang, M.H.: Object tracking benchmark. IEEE Trans. Pattern Anal. Mach. Intell. 37, 1834\u20131848 (2015)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."}],"container-title":["The Visual Computer"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00371-021-02150-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00371-021-02150-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00371-021-02150-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,7,19]],"date-time":"2022-07-19T09:07:47Z","timestamp":1658221667000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00371-021-02150-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,5,13]]},"references-count":36,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2022,8]]}},"alternative-id":["2150"],"URL":"https:\/\/doi.org\/10.1007\/s00371-021-02150-1","relation":{},"ISSN":["0178-2789","1432-2315"],"issn-type":[{"value":"0178-2789","type":"print"},{"value":"1432-2315","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,5,13]]},"assertion":[{"value":"23 April 2021","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 May 2021","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}