{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T06:13:22Z","timestamp":1774419202880,"version":"3.50.1"},"reference-count":101,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2020,1,19]],"date-time":"2020-01-19T00:00:00Z","timestamp":1579392000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Air Force Research Laboratory, Sensors Directorate (AFRL\/RYAP) contract to Systems and Technology Research","award":["FA8650-18-C-1739"],"award-info":[{"award-number":["FA8650-18-C-1739"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In recent years, deep learning-based visual object trackers have achieved state-of-the-art performance on several visual object tracking benchmarks. However, most tracking benchmarks are focused on ground level videos, whereas aerial tracking presents a new set of challenges. In this paper, we compare ten trackers based on deep learning techniques on four aerial datasets. We choose top performing trackers utilizing different approaches, specifically tracking by detection, discriminative correlation filters, Siamese networks and reinforcement learning. In our experiments, we use a subset of OTB2015 dataset with aerial style videos; the UAV123 dataset without synthetic sequences; the UAV20L dataset, which contains 20 long sequences; and DTB70 dataset as our benchmark datasets. We compare the advantages and disadvantages of different trackers in different tracking situations encountered in aerial data. Our findings indicate that the trackers perform significantly worse in aerial datasets compared to standard ground level videos. We attribute this effect to smaller target size, camera motion, significant camera rotation with respect to the target, out of view movement, and clutter in the form of occlusions or similar looking distractors near tracked object.<\/jats:p>","DOI":"10.3390\/s20020547","type":"journal-article","created":{"date-parts":[[2020,1,21]],"date-time":"2020-01-21T03:04:43Z","timestamp":1579575883000},"page":"547","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["Benchmarking Deep Trackers on Aerial Videos"],"prefix":"10.3390","volume":"20","author":[{"given":"Abu Md Niamul","family":"Taufique","sequence":"first","affiliation":[{"name":"Rochester Institute of Technology, Rochester, NY 14623, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Breton","family":"Minnehan","sequence":"additional","affiliation":[{"name":"Rochester Institute of Technology, Rochester, NY 14623, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andreas","family":"Savakis","sequence":"additional","affiliation":[{"name":"Rochester Institute of Technology, Rochester, NY 14623, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,1,19]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1109\/70.210792","article-title":"Visual tracking of a moving target by a camera mounted on a robot: A combination of control and vision","volume":"9","author":"Papanikolopoulos","year":"1993","journal-title":"IEEE Trans. Robot. Autom."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1429","DOI":"10.1109\/TMM.2015.2455418","article-title":"On-Road Pedestrian Tracking Across Multiple Driving Recorders","volume":"17","author":"Lee","year":"2015","journal-title":"IEEE Trans. Multimed."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Laurense, V.A., Goh, J.Y., and Gerdes, J.C. (2017, January 24\u201326). Path-tracking for autonomous vehicles at the limit of friction. Proceedings of the 2017 American Control Conference (ACC), Seattle, WA, USA.","DOI":"10.23919\/ACC.2017.7963824"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Tang, S., Andriluka, M., Andres, B., and Schiele, B. (2017, January 21\u201326). Multiple people tracking by lifted multicut and person reidentification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.394"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Girdhar, R., Gkioxari, G., Torresani, L., Paluri, M., and Tran, D. (2018, January 18\u201322). Detect-and-track: Efficient pose estimation in videos. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00044"},{"key":"ref_6","unstructured":"Walker, S., Sewell, C., Park, J., Ravindran, P., Koolwal, A., Camarillo, D., and Barbagli, F. (2017). Systems and Methods for Localizing, Tracking and\/or Controlling Medical Instruments. (App. 15\/466,565), U.S. Patent."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Speidel, S., Kuhn, E., Bodenstedt, S., R\u00f6hl, S., Kenngott, H., M\u00fcller-Stich, B., and Dillmann, R. (2014, January 12). Visual tracking of da vinci instruments for laparoscopic surgery. Proceedings of the SPIE Medical Imaging 2014: Image-Guided Procedures, Robotic Interventions, and Modeling, San Diego, CA, USA.","DOI":"10.1117\/12.2042483"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"70","DOI":"10.1016\/j.patrec.2014.04.011","article-title":"Human activity recognition from 3d data: A review","volume":"48","author":"Aggarwal","year":"2014","journal-title":"Pattern Recognit. Lett."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1145\/1177352.1177355","article-title":"Object tracking: A survey","volume":"38","author":"Yilmaz","year":"2006","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1442","DOI":"10.1109\/TPAMI.2013.230","article-title":"Visual tracking: An experimental survey","volume":"36","author":"Smeulders","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1245","DOI":"10.1142\/S0218001409007624","article-title":"Robust object tracking using joint color-texture histogram","volume":"23","author":"Ning","year":"2009","journal-title":"Int. J. Pattern Recognit. Artif. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1016\/j.cviu.2008.08.006","article-title":"Object tracking using SIFT features and mean shift","volume":"113","author":"Zhou","year":"2009","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Bolme, D.S., Beveridge, J.R., Draper, B.A., and Lui, Y.M. (2010, January 13\u201318). Visual object tracking using adaptive correlation filters. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539960"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"2170","DOI":"10.1016\/j.patcog.2011.03.002","article-title":"Visual object tracking via sample-based Adaptive Sparse Representation (AdaSR)","volume":"44","author":"Han","year":"2011","journal-title":"Pattern Recognit."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"2340","DOI":"10.1109\/TIP.2011.2174370","article-title":"Particle filter with a mode tracker for visual tracking across illumination changes","volume":"21","author":"Das","year":"2011","journal-title":"IEEE Trans. Image Process."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Danelljan, M., H\u00e4ger, G., Khan, F., and Felsberg, M. (2014, January 1\u20135). Accurate scale estimation for robust visual tracking. Proceedings of the British Machine Vision Conference, Nottingham, UK.","DOI":"10.5244\/C.28.65"},{"key":"ref_17","unstructured":"Jia, X., Lu, H., and Yang, M.H. (2012, January 16\u201321). Visual tracking via adaptive structural local sparse appearance model. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA."},{"key":"ref_18","unstructured":"Zhong, W., Lu, H., and Yang, M.H. (2012, January 16\u201321). Robust object tracking via sparsity-based collaborative model. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yoon, J.H., Kim, D.Y., and Yoon, K.J. (2012). Visual tracking via adaptive tracker selection with multiple features. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-642-33765-9_3"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"583","DOI":"10.1109\/TPAMI.2014.2345390","article-title":"High-speed tracking with kernelized correlation filters","volume":"37","author":"Henriques","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2096","DOI":"10.1109\/TPAMI.2015.2509974","article-title":"Struck: Structured output tracking with kernels","volume":"38","author":"Hare","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Kiani Galoogahi, H., Fagg, A., and Lucey, S. (2017, January 22\u201329). Learning background-aware correlation filters for visual tracking. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.129"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Danelljan, M., Hager, G., Shahbaz Khan, F., and Felsberg, M. (2015, January 7\u201313). Learning spatially regularized correlation filters for visual tracking. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.490"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"013006","DOI":"10.1117\/1.JEI.24.1.013006","article-title":"Visual tracking with multifeature joint sparse representation","volume":"24","author":"Dong","year":"2015","journal-title":"J. Electron. Imaging"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Ontiveros-Gallardo, S.E., and Kober, V. (2015, January 9). Objects tracking with adaptive correlation filters and kalman filtering. Proceedings of the SPIE Optics and Photonics for Information Processing IX, San Diego, CA, USA.","DOI":"10.1117\/12.2187109"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Leal-Taix\u00e9, L., and Roth, S. (2019). The Sixth Visual Object Tracking VOT2018 Challenge Results. European Conference in Computer Vision ECCV 2018 Workshops, Springer.","DOI":"10.1007\/978-3-030-11024-6"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1834","DOI":"10.1109\/TPAMI.2014.2388226","article-title":"Object tracking benchmark","volume":"37","author":"Wu","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Danelljan, M., Bhat, G., Khan, F.S., and Felsberg, M. (2017, January 21\u201326). ECO: Efficient Convolution Operators for Tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.733"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Danelljan, M., Robinson, A., Khan, F.S., and Felsberg, M. (2016). Beyond correlation filters: Learning continuous convolution operators for visual tracking. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46454-1_29"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Bertinetto, L., Valmadre, J., Henriques, J.F., Vedaldi, A., and Torr, P.H. (2016). Fully-convolutional Siamese networks for object tracking. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-48881-3_56"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Zhu, Z., Wang, Q., Li, B., Wu, W., Yan, J., and Hu, W. (2018). Distractor-aware Siamese Networks for Visual Object Tracking. arXiv.","DOI":"10.1007\/978-3-030-01240-3_7"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Li, B., Yan, J., Wu, W., Zhu, Z., and Hu, X. (2018, January 18\u201322). High Performance Visual Tracking With Siamese Region Proposal Network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00935"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Nam, H., and Han, B. (2016, January 27\u201330). Learning multi-domain convolutional neural networks for visual tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.465"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Jung, I., Son, J., Baek, M., and Han, B. (2018, January 8\u201314). Real-time mdnet. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01225-0_6"},{"key":"ref_35","unstructured":"Pu, S., Song, Y., Ma, C., Zhang, H., and Yang, M.H. (2018). Deep Attentive Tracking via Reciprocative Learning. Neural Information Processing Systems, NIPS."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Yun, S., Choi, J., Yoo, Y., Yun, K., and Young Choi, J. (2017, January 21\u201326). Action-decision networks for visual tracking with deep reinforcement learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.148"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Chen, B., Wang, D., Li, P., Wang, S., and Lu, H. (2018, January 8\u201314). Real-time\u2019Actor-Critic\u2019Tracking. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_20"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Wu, Y., Lim, J., and Yang, M.H. (2013, January 23\u201328). Online Object Tracking: A Benchmark. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.312"},{"key":"ref_39","unstructured":"Kristan, M., Leonardis, A., Matas, J., Felsberg, M., Pflugfelder, R., Cehovin Zajc, L., Vojir, T., Hager, G., Lukezic, A., and Eldesokey, A. (2017, January 2\u201329). The visual object tracking VOT2017 challenge results. Proceedings of the ICCV2017 Workshops, Workshop on Visual Object Tracking Challenge, Venice, Italy."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Hua, G., and J\u00e9gou, H. (2016). The Visual Object Tracking VOT2016 Challenge Results. European Conference in Computer Vision\u2014ECCV 2016 Workshops, Springer.","DOI":"10.1007\/978-3-319-48881-3"},{"key":"ref_41","unstructured":"Kristan, M., Matas, J., Leonardis, A., Felsberg, M., Cehovin, L., Fernandez, G., Vojir, T., Hager, G., Nebehay, G., and Pflugfelder, R. (2015, January 11\u201318). The visual object tracking VOT2015 challenge results. Proceedings of the ICCV2015 Workshops, Workshop on Visual Object Tracking Challenge, Santiago, Chile."},{"key":"ref_42","unstructured":"Kristan, M., Pflugfelder, R., Leonardis, A., Matas, J., Cehovin, L., Nebehay, G., Gustavo, F., Vojir, T., Dimitriev, A., and Petrosino, A. (2014, January 6\u201312). The visual object tracking VOT2014 challenge results. Proceedings of the ECCV2014 Workshops, Workshop on Visual Object Tracking Challenge, Zurich, Switzerland."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Kristan, M., Pflugfelder, R., Leonardis, A., Matas, J., Porikli, F., Cehovin, L., Nebehay, G., and Vojir, T. (2013, January 1\u20138). The visual object tracking VOT2013 challenge results. Proceedings of the ICCV2013 Workshops, Workshop on Visual Object Tracking Challenge, Sydney, Australia.","DOI":"10.1109\/ICCVW.2013.20"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Song, S., and Xiao, J. (2013, January 1\u20138). Tracking revisited using RGBD camera: Unified benchmark and baselines. Proceedings of the IEEE International Conference on Computer Vision, Sydney, Australia.","DOI":"10.1109\/ICCV.2013.36"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1109\/TPAMI.2015.2417577","article-title":"NUS-PRO: A new visual tracking challenge","volume":"38","author":"Li","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"5630","DOI":"10.1109\/TIP.2015.2482905","article-title":"Encoding color information for visual tracking: Algorithms and benchmark","volume":"24","author":"Liang","year":"2015","journal-title":"IEEE Trans. Image Process."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Mueller, M., Smith, N., and Ghanem, B. (2016). A benchmark and simulator for UAV tracking. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46448-0_27"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Li, S., and Yeung, D.Y. (2017). Visual Object Tracking for Unmanned Aerial Vehicles: A Benchmark and New Motion Models, AAAI.","DOI":"10.1609\/aaai.v31i1.11205"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Kiani Galoogahi, H., Fagg, A., Huang, C., Ramanan, D., and Lucey, S. (2017, January 22\u201329). Need for speed: A benchmark for higher frame rate object tracking. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.128"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Zajc, L.C., Luke\u017eic, A., Leonardis, A., and Kristan, M. (2017, January 22\u201329). Beyond standard benchmarks: Parameterizing performance evaluation in visual object tracking. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.360"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Du, D., Qi, Y., Yu, H., Yang, Y., Duan, K., Li, G., Zhang, W., Huang, Q., and Tian, Q. (2018, January 8\u201314). The unmanned aerial vehicle benchmark: Object detection and tracking. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01249-6_23"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Muller, M., Bibi, A., Giancola, S., Alsubaihi, S., and Ghanem, B. (2018, January 8\u201314). Trackingnet: A large-scale dataset and benchmark for object tracking in the wild. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01246-5_19"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Valmadre, J., Bertinetto, L., Henriques, J.F., Tao, R., Vedaldi, A., Smeulders, A.W., Torr, P.H., and Gavves, E. (2018, January 8\u201314). Long-term tracking in the wild: A benchmark. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01219-9_41"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Fan, H., Lin, L., Yang, F., Chu, P., Deng, G., Yu, S., Bai, H., Xu, Y., Liao, C., and Ling, H. (2018). LaSOT: A high-quality benchmark for large-scale single object tracking. arXiv.","DOI":"10.1109\/CVPR.2019.00552"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Minnehan, B., Salmin, A., Salva, K., and Savakis, A. (2018, January 30). Benchmarking deep learning trackers on aerial videos. Proceedings of the SPIE Pattern Recognition and Tracking XXIX, Orlando, FL, USA.","DOI":"10.1117\/12.2323866"},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"323","DOI":"10.1016\/j.patcog.2017.11.007","article-title":"Deep visual tracking: Review and experimental comparison","volume":"76","author":"Li","year":"2018","journal-title":"Pattern Recognit."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Fiaz, M., Mahmood, A., Javed, S., and Jung, S.K. (2018). Handcrafted and Deep Trackers: A Review of Recent Object Tracking Approaches. arXiv.","DOI":"10.1145\/3309665"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Song, Y., Ma, C., Wu, X., Gong, L., Bao, L., Zuo, W., Shen, C., Lau, R.W., and Yang, M.H. (2018, January 18\u201322). Vital: Visual tracking via adversarial learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00937"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Park, E., and Berg, A.C. (2018). Meta-Tracker: Fast and Robust Online Adaptation for Visual Object Trackers. arXiv.","DOI":"10.1007\/978-3-030-01219-9_35"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Fan, H., and Ling, H. (2017, January 21\u201326). Sanet: Structure-aware network for visual tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.275"},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Teng, Z., Xing, J., Wang, Q., Lang, C., Feng, S., and Jin, Y. (2017, January 22\u201329). Robust object tracking based on temporal and spatial deep networks. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.130"},{"key":"ref_62","unstructured":"Nam, H., Baek, M., and Han, B. (2016). Modeling and propagating cnns in a tree structure for visual tracking. arXiv."},{"key":"ref_63","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"ImageNet large scale visual recognition challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Valmadre, J., Bertinetto, L., Henriques, J., Vedaldi, A., and Torr, P.H. (2017, January 21\u201326). End-to-end representation learning for correlation filter based tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.531"},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Kart, U., Lukezic, A., Kristan, M., Kamarainen, J.K., and Matas, J. (2018). Object Tracking by Reconstruction with View-Specific Discriminative Correlation Filters. arXiv.","DOI":"10.1109\/CVPR.2019.00143"},{"key":"ref_67","doi-asserted-by":"crossref","first-page":"2526","DOI":"10.1109\/TIP.2018.2806280","article-title":"Good features to correlate for visual tracking","volume":"27","author":"Gundogdu","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Bhat, G., Johnander, J., Danelljan, M., Shahbaz Khan, F., and Felsberg, M. (2018, January 8\u201314). Unveiling the power of deep tracking. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01216-8_30"},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Zhang, M., Wang, Q., Xing, J., Gao, J., Peng, P., Hu, W., and Maybank, S. (2018, January 8\u201314). Visual tracking via spatially aligned correlation filters network. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01219-9_29"},{"key":"ref_70","doi-asserted-by":"crossref","unstructured":"Zhu, Z., Wu, W., Zou, W., and Yan, J. (2018, January 18\u201322). End-to-end flow correlation tracking with spatial-temporal attention. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00064"},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Yao, Y., Wu, X., Zhang, L., Shan, S., and Zuo, W. (2018, January 8\u201314). Joint representation and truncated inference learning for correlation filter based tracking. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01240-3_34"},{"key":"ref_72","doi-asserted-by":"crossref","unstructured":"Wang, N., Zhou, W., Tian, Q., Hong, R., Wang, M., and Li, H. (2018, January 18\u201322). Multi-cue correlation filters for robust visual tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00509"},{"key":"ref_73","doi-asserted-by":"crossref","unstructured":"Tang, M., Yu, B., Zhang, F., and Wang, J. (2018, January 18\u201322). High-speed tracking with multi-kernel correlation filters. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00512"},{"key":"ref_74","doi-asserted-by":"crossref","unstructured":"He, Z., Fan, Y., Zhuang, J., Dong, Y., and Bai, H. (2017, January 22\u201329). Correlation filters with weighted convolution responses. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCVW.2017.233"},{"key":"ref_75","doi-asserted-by":"crossref","unstructured":"Li, F., Yao, Y., Li, P., Zhang, D., Zuo, W., and Yang, M.H. (2017, January 22\u201329). Integrating boundary and center correlation filters for visual tracking with aspect ratio variation. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCVW.2017.234"},{"key":"ref_76","doi-asserted-by":"crossref","unstructured":"Mueller, M., Smith, N., and Ghanem, B. (2017, January 21\u201326). Context-aware correlation filter tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.152"},{"key":"ref_77","doi-asserted-by":"crossref","unstructured":"Zhang, T., Xu, C., and Yang, M.H. (2017, January 21\u201326). Multi-task correlation particle filter for robust object tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.512"},{"key":"ref_78","doi-asserted-by":"crossref","unstructured":"Choi, J., Jin Chang, H., Yun, S., Fischer, T., Demiris, Y., and Young Choi, J. (2017, January 21\u201326). Attentional correlation filter network for adaptive visual tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.513"},{"key":"ref_79","doi-asserted-by":"crossref","unstructured":"Bibi, A., Mueller, M., and Ghanem, B. (2016). Target response adaptation for correlation filter tracking. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46466-4_25"},{"key":"ref_80","doi-asserted-by":"crossref","unstructured":"Li, F., Tian, C., Zuo, W., Zhang, L., and Yang, M.H. (2018). Learning Spatial-Temporal Regularized Correlation Filters for Visual Tracking. arXiv.","DOI":"10.1109\/CVPR.2018.00515"},{"key":"ref_81","unstructured":"Bai, S., He, Z., Xu, T.B., Zhu, Z., Dong, Y., and Bai, H. (2018). Multi-hierarchical Independent Correlation Filters for Visual Tracking. arXiv."},{"key":"ref_82","first-page":"551","article-title":"Online passive-aggressive algorithms","volume":"7","author":"Crammer","year":"2006","journal-title":"J. Mach. Learn. Res."},{"key":"ref_83","doi-asserted-by":"crossref","unstructured":"Wang, Q., Zhang, L., Bertinetto, L., Hu, W., and Torr, P.H. (2019, January 24\u201327). Fast Online Object Tracking and Segmentation: A Unifying Approach. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2019.00142"},{"key":"ref_84","doi-asserted-by":"crossref","unstructured":"Li, B., Wu, W., Wang, Q., Zhang, F., Xing, J., and Yan, J. (2018). SiamRPN++: Evolution of Siamese Visual Tracking with Very Deep Networks. arXiv.","DOI":"10.1109\/CVPR.2019.00441"},{"key":"ref_85","doi-asserted-by":"crossref","unstructured":"Zhang, Z., and Peng, H. (2019, January 24\u201327). Deeper and Wider Siamese Networks for Real-Time Visual Tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2019.00472"},{"key":"ref_86","doi-asserted-by":"crossref","unstructured":"Li, X., Ma, C., Wu, B., He, Z., and Yang, M.H. (2019, January 24\u201327). Target-Aware Deep Tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2019.00146"},{"key":"ref_87","doi-asserted-by":"crossref","unstructured":"Fan, H., and Ling, H. (2018). Siamese Cascaded Region Proposal Networks for Real-Time Visual Tracking. arXiv.","DOI":"10.1109\/CVPR.2019.00814"},{"key":"ref_88","doi-asserted-by":"crossref","unstructured":"Wang, G., Luo, C., Xiong, Z., and Zeng, W. (2019). SPM-Tracker: Series-Parallel Matching for Real-Time Visual Object Tracking. arXiv.","DOI":"10.1109\/CVPR.2019.00376"},{"key":"ref_89","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Wang, L., Qi, J., Wang, D., Feng, M., and Lu, H. (2018, January 8\u201314). Structured Siamese network for real-time visual tracking. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01240-3_22"},{"key":"ref_90","doi-asserted-by":"crossref","unstructured":"Dong, X., and Shen, J. (2018, January 8\u201314). Triplet loss in Siamese network for object tracking. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01261-8_28"},{"key":"ref_91","doi-asserted-by":"crossref","unstructured":"Wang, Q., Teng, Z., Xing, J., Gao, J., Hu, W., and Maybank, S. (2018, January 18\u201322). Learning attentions: Residual attentional Siamese network for high performance online visual tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00510"},{"key":"ref_92","doi-asserted-by":"crossref","unstructured":"Guo, Q., Feng, W., Zhou, C., Huang, R., Wan, L., and Wang, S. (2017, January 22\u201329). Learning dynamic Siamese network for visual object tracking. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.196"},{"key":"ref_93","doi-asserted-by":"crossref","unstructured":"Minnehan, B., Taufique, A.M.N., and Savakis, A. (2019, January 13). Fully convolutional adaptive tracker with real time performance. Proceedings of the SPIE Geospatial Informatics IX, Baltimore, MD, USA.","DOI":"10.1117\/12.2518823"},{"key":"ref_94","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015). Faster R-CNN: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems, NIPS."},{"key":"ref_95","doi-asserted-by":"crossref","unstructured":"He, A., Luo, C., Tian, X., and Zeng, W. (2018, January 18\u201322). A twofold Siamese network for real-time object tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00508"},{"key":"ref_96","doi-asserted-by":"crossref","unstructured":"Real, E., Shlens, J., Mazzocchi, S., Pan, X., and Vanhoucke, V. (2017, January 21\u201326). Youtube-boundingboxes: A large high-precision human-annotated data set for object detection in video. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.789"},{"key":"ref_97","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014). Microsoft coco: Common objects in context. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_98","doi-asserted-by":"crossref","unstructured":"Ren, L., Yuan, X., Lu, J., Yang, M., and Zhou, J. (2018, January 8\u201314). Deep reinforcement learning with iterative shift for visual tracking. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01240-3_42"},{"key":"ref_99","doi-asserted-by":"crossref","unstructured":"Dong, X., Shen, J., Wang, W., Liu, Y., Shao, L., and Porikli, F. (2018, January 18\u201322). Hyperparameter optimization for tracking with continuous deep q-learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00061"},{"key":"ref_100","doi-asserted-by":"crossref","unstructured":"Supancic, J., and Ramanan, D. (2017, January 22\u201329). Tracking as online decision-making: Learning a policy from streaming videos with reinforcement learning. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.43"},{"key":"ref_101","doi-asserted-by":"crossref","unstructured":"Huang, C., Lucey, S., and Ramanan, D. (2017, January 22\u201329). Learning policies for adaptive tracking with deep feature cascades. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.21"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/2\/547\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T14:05:26Z","timestamp":1760364326000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/2\/547"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,1,19]]},"references-count":101,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2020,1]]}},"alternative-id":["s20020547"],"URL":"https:\/\/doi.org\/10.3390\/s20020547","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,1,19]]}}}