{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,16]],"date-time":"2026-01-16T11:15:48Z","timestamp":1768562148912,"version":"3.49.0"},"reference-count":33,"publisher":"Springer Science and Business Media LLC","issue":"32","license":[{"start":{"date-parts":[[2024,8,12]],"date-time":"2024-08-12T00:00:00Z","timestamp":1723420800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,8,12]],"date-time":"2024-08-12T00:00:00Z","timestamp":1723420800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100010661","name":"Horizon 2020 Framework Programme","doi-asserted-by":"publisher","award":["871449"],"award-info":[{"award-number":["871449"]}],"id":[{"id":"10.13039\/100010661","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007605","name":"Aarhus Universitet","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100007605","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2024,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In this paper, we propose a novel voxel-based 3D single object tracking (3D SOT) method called Voxel Pseudo Image Tracking (VPIT). VPIT is the first method that uses voxel pseudo images for 3D SOT. The input point cloud is structured by pillar-based voxelization, and the resulting pseudo image is used as an input to a 2D-like Siamese SOT method. The pseudo image is created in the Bird\u2019s-eye View (BEV) coordinates; and therefore, the objects in it have constant size. Thus, only the object rotation can change in the new coordinate system and not the object scale. For this reason, we replace multi-scale search with a multi-rotation search, where differently rotated search regions are compared against a single target representation to predict both position and rotation of the object. Experiments on KITTI [1] Tracking dataset show that VPIT is the fastest 3D SOT method and maintains competitive Success and Precision values. Application of a SOT method in a real-world scenario meets with limitations such as lower computational capabilities of embedded devices and a latency-unforgiving environment, where the method is forced to skip certain data frames if the inference speed is not high enough. We implement a real-time evaluation protocol and show that other methods lose most of their performance on embedded devices; while, VPIT maintains its ability to track the object.<\/jats:p>","DOI":"10.1007\/s00521-024-10259-2","type":"journal-article","created":{"date-parts":[[2024,8,12]],"date-time":"2024-08-12T14:03:07Z","timestamp":1723471387000},"page":"20341-20354","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Vpit: real-time embedded single object 3D tracking using voxel pseudo images"],"prefix":"10.1007","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7592-365X","authenticated-orcid":false,"given":"Illia","family":"Oleksiienko","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paraskevi","family":"Nousi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nikolaos","family":"Passalis","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anastasios","family":"Tefas","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alexandros","family":"Iosifidis","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,8,12]]},"reference":[{"key":"10259_CR1","doi-asserted-by":"crossref","unstructured":"Geiger A, Lenz P, Urtasun R (2012) Are we ready for autonomous driving? The KITTI vision benchmark suite. In: IEEE conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"10259_CR2","unstructured":"Geiger A, Lenz P, Stiller C, Urtasun R (2023) KITTI 3D object detection leaderboard. [Accessed 15-December-2023]. http:\/\/www.cvlibs.net\/datasets\/kitti\/eval_object.php?obj_benchmark=3d"},{"key":"10259_CR3","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2020.113816","volume":"165","author":"C Badue","year":"2021","unstructured":"Badue C, Guidolini R, Carneiro RV, Azevedo P, Cardoso VB, Forechi A, Jesus L, Berriel R, Paix\u00e3o TM, Mutz F, de Paula Veronese L, Oliveira-Santos T, De Souza AF (2021) Self-driving cars: a survey. Expert Syst Appl 165:113816","journal-title":"Expert Syst Appl"},{"key":"10259_CR4","doi-asserted-by":"crossref","unstructured":"Bolme DS, Beveridge JR, Draper BA, Lui YM (2010) Visual object tracking using adaptive correlation filters. In: 2010 IEEE computer society conference on computer vision and pattern recognition, pp 2544\u20132550","DOI":"10.1109\/CVPR.2010.5539960"},{"issue":"3","key":"10259_CR5","doi-asserted-by":"publisher","first-page":"583","DOI":"10.1109\/TPAMI.2014.2345390","volume":"37","author":"JF Henriques","year":"2015","unstructured":"Henriques JF, Caseiro R, Martins P, Batista J (2015) High-speed tracking with kernelized correlation filters. IEEE Trans Pattern Anal Mach Intell 37(3):583\u2013596","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"10259_CR6","doi-asserted-by":"crossref","unstructured":"Held D, Thrun S, Savarese S (2016) Learning to track at 100 fps with deep regression networks. 1604.01802","DOI":"10.1007\/978-3-319-46448-0_45"},{"issue":"4","key":"10259_CR7","doi-asserted-by":"publisher","first-page":"4995","DOI":"10.1109\/JSEN.2020.3033034","volume":"21","author":"Z Fang","year":"2021","unstructured":"Fang Z, Zhou S, Cui Y, Scherer S (2021) 3d-siamrpn: an end-to-end learning method for real-time 3d single object tracking using raw point cloud. IEEE Sens J 21(4):4995\u20135011","journal-title":"IEEE Sens J"},{"key":"10259_CR8","doi-asserted-by":"crossref","unstructured":"Bertinetto L, Valmadre J, Henriques JF, Vedaldi A, Torr PH (2016) Fully-convolutional siamese networks for object tracking. arXiv:1606.09549","DOI":"10.1007\/978-3-319-48881-3_56"},{"key":"10259_CR9","doi-asserted-by":"crossref","unstructured":"Li B, Yan J, Wu W, Zhu Z, Hu X (2018) High performance visual tracking with siamese region proposal network. In: proceedings of the IEEE conference on computer vision and pattern recognition, pp 8971\u20138980","DOI":"10.1109\/CVPR.2018.00935"},{"key":"10259_CR10","doi-asserted-by":"crossref","unstructured":"Li B, Wu W, Wang Q, Zhang F, Xing J, Yan J (2018) Siamrpn++: evolution of siamese visual tracking with very deep networks. arXiv:1812.11703","DOI":"10.1109\/CVPR.2019.00441"},{"key":"10259_CR11","doi-asserted-by":"crossref","unstructured":"Qi H, Feng C, Cao Z, Zhao F, Xiao Y (2020) P2b: point-to-box network for 3d object tracking in point clouds. arXiv:2005.13888","DOI":"10.1109\/CVPR42600.2020.00636"},{"key":"10259_CR12","doi-asserted-by":"publisher","first-page":"2339","DOI":"10.1109\/TMM.2022.3146714","volume":"25","author":"J Shan","year":"2023","unstructured":"Shan J, Zhou S, Cui Y, Fang Z (2023) Real-time 3d single object tracking with transformer. IEEE Trans Multimed 25:2339\u20132353","journal-title":"IEEE Trans Multimed"},{"key":"10259_CR13","doi-asserted-by":"crossref","unstructured":"Li M, Wang Y-X, Ramanan D (2020) Towards streaming perception. In: European conference on computer vision, pp 473\u2013488. Springer","DOI":"10.1007\/978-3-030-58536-5_28"},{"issue":"8","key":"10259_CR14","doi-asserted-by":"publisher","first-page":"3412","DOI":"10.1109\/TNNLS.2020.3015992","volume":"32","author":"Y Li","year":"2021","unstructured":"Li Y, Ma L, Zhong Z, Liu F, Chapman MA, Cao D, Li J (2021) Deep learning for lidar point clouds in autonomous driving: a review. IEEE Trans Neural Netw Learn Syst 32(8):3412\u20133432","journal-title":"IEEE Trans Neural Netw Learn Syst"},{"key":"10259_CR15","doi-asserted-by":"crossref","unstructured":"Giancola S, Zarzar J, Ghanem B (2019) Leveraging shape completion for 3D siamese tracking. In: proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR), pp 1359\u20131368","DOI":"10.1109\/CVPR.2019.00145"},{"key":"10259_CR16","doi-asserted-by":"crossref","unstructured":"Zheng C, Yan X, Gao J, Zhao W, Zhang W, Li Z, Cui S (2021) Box-aware feature enhancement for single object tracking on point clouds. In: proceedings of the IEEE\/CVF international conference on computer vision (CVPR), pp 13199\u201313208","DOI":"10.1109\/ICCV48922.2021.01295"},{"key":"10259_CR17","doi-asserted-by":"crossref","unstructured":"Zou H, Cui J, Kong X, Zhang C, Liu Y, Wen F, Li W (2020) F-siamese tracker: A frustum-based double siamese network for 3d single object tracking. In: 2020 IEEE\/RSJ international conference on intelligent robots and systems (IROS), pp 8133\u20138139. IEEE","DOI":"10.1109\/IROS45743.2020.9341120"},{"key":"10259_CR18","doi-asserted-by":"crossref","unstructured":"Hui L, Wang L, Tang L, Lan K, Xie J, Yang J (2022) 3D siamese transformer network for single object tracking on point clouds. In: proceedings of the European Conference on Computer Vision (ECCV)","DOI":"10.1007\/978-3-031-20086-1_17"},{"key":"10259_CR19","doi-asserted-by":"crossref","unstructured":"Zhou C, Luo Z, Luo Y, Liu T, Pan L, Cai Z, Zhao H, Lu S (2022) Pttr: Relational 3d point cloud object tracking with transformer. In: proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR)","DOI":"10.1109\/CVPR52688.2022.00834"},{"key":"10259_CR20","unstructured":"Qi CR, Yi L, Su H, Guibas LJ (2017) Pointnet++: deep hierarchical feature learning on point sets in a metric space. arXiv:1706.02413"},{"key":"10259_CR21","doi-asserted-by":"crossref","unstructured":"Nie J, He Z, Yang Y, Gao M, Zhang J (2023) Glt-t: Global-local transformer voting for 3d single object tracking in point clouds. In: proceedings of the AAAI conference on artificial intelligence, vol. 37, pp 1957\u20131965","DOI":"10.1609\/aaai.v37i2.25287"},{"key":"10259_CR22","doi-asserted-by":"crossref","unstructured":"Qi CR, Litany O, He K, Guibas LJ (2019) Deep hough voting for 3d object detection in point clouds. In: proceedings of the IEEE international conference on computer vision","DOI":"10.1109\/ICCV.2019.00937"},{"key":"10259_CR23","unstructured":"Zarzar J, Giancola S, Ghanem B (2020) Efficient bird eye view proposals for 3d siamese tracking. arXiv:1903.10168"},{"key":"10259_CR24","doi-asserted-by":"crossref","unstructured":"Kristan M, Matas J, Leonardis A, Felsberg M, Pflugfelder R, K\u00e4m\u00e4r\u00e4inen J-K, Chang HJ, Danelljan M, Cehovin L, Luke\u017ei\u010d A etal: (2021) The ninth visual object tracking vot2021 challenge results. In: proceedings of the IEEE\/CVF international conference on computer vision, pp 2711\u20132738","DOI":"10.1109\/ICCVW54120.2021.00305"},{"key":"10259_CR25","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2020.103933","volume":"99","author":"P Nousi","year":"2020","unstructured":"Nousi P, Tefas A, Pitas I (2020) Dense convolutional feature histograms for robust visual object tracking. Image Vis Comput 99:103933","journal-title":"Image Vis Comput"},{"key":"10259_CR26","doi-asserted-by":"crossref","unstructured":"Zhou Y, Tuzel O (2018) VoxelNet: end-to-end Learning for Point Cloud Based 3D Object Detection. In: IEEE conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR.2018.00472"},{"key":"10259_CR27","doi-asserted-by":"crossref","unstructured":"Lang AH, Vora S, Caesar H, Zhou L, Yang J, Beijbom O (2019) PointPillars: fast encoders for object detection from point clouds. In: IEEE conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR.2019.01298"},{"key":"10259_CR28","doi-asserted-by":"crossref","unstructured":"Liu Z, Zhao X, Huang T, Hu R, Zhou Y, Bai X (2020) TANet: robust 3D object detection from point clouds with triple attention. In: AAAI Conference on Artificial Intelligence","DOI":"10.1609\/aaai.v34i07.6837"},{"key":"10259_CR29","doi-asserted-by":"crossref","unstructured":"Chen Q, Sun L, Wang Z, Jia K, Yuille AL (2020) Object as Hotspots: an anchor-free 3D object detection approach via firing of hotspots. In: European conference on computer vision","DOI":"10.1007\/978-3-030-58589-1_5"},{"key":"10259_CR30","unstructured":"Qi CR, Su H, Mo K, Guibas LJ (2017) Pointnet: deep learning on point sets for 3d classification and segmentation. arXiv:1612.00593"},{"key":"10259_CR31","doi-asserted-by":"crossref","unstructured":"Caesar H, Bankiti V, Lang AH, Vora S, Liong VE, Xu Q, Krishnan A, Pan Y, Baldan G, Beijbom O (2020) nuScenes: a multimodal dataset for autonomous driving. In: IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR42600.2020.01164"},{"issue":"11","key":"10259_CR32","doi-asserted-by":"publisher","first-page":"2137","DOI":"10.1109\/TPAMI.2016.2516982","volume":"38","author":"M Kristan","year":"2016","unstructured":"Kristan M, Matas J, Leonardis A, Vojir T, Pflugfelder R, Fernandez G, Nebehay G, Porikli F, Cehovin L (2016) A novel performance evaluation methodology for single-target trackers. IEEE Trans Pattern Anal Mach Intell 38(11):2137\u20132155","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"1","key":"10259_CR33","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1115\/1.3662552","volume":"82","author":"RE Kalman","year":"1960","unstructured":"Kalman RE (1960) A New Approach to Linear Filtering and Prediction Problems. J Basic Eng 82(1):35\u201345","journal-title":"J Basic Eng"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-024-10259-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-024-10259-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-024-10259-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,28]],"date-time":"2024-09-28T08:08:24Z","timestamp":1727510904000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-024-10259-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,12]]},"references-count":33,"journal-issue":{"issue":"32","published-print":{"date-parts":[[2024,11]]}},"alternative-id":["10259"],"URL":"https:\/\/doi.org\/10.1007\/s00521-024-10259-2","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"value":"0941-0643","type":"print"},{"value":"1433-3058","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,12]]},"assertion":[{"value":"18 December 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 July 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 August 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}