{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,26]],"date-time":"2026-03-26T09:37:44Z","timestamp":1774517864395,"version":"3.50.1"},"reference-count":33,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2021,4,21]],"date-time":"2021-04-21T00:00:00Z","timestamp":1618963200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Multi-Object Tracking (MOT) is an integral part of any autonomous driving pipelines because it produces trajectories of other moving objects in the scene and predicts their future motion. Thanks to the recent advances in 3D object detection enabled by deep learning, track-by-detection has become the dominant paradigm in 3D MOT. In this paradigm, a MOT system is essentially made of an object detector and a data association algorithm which establishes track-to-detection correspondence. While 3D object detection has been actively researched, association algorithms for 3D MOT has settled at bipartite matching formulated as a Linear Assignment Problem (LAP) and solved by the Hungarian algorithm. In this paper, we adapt a two-stage data association method which was successfully applied to image-based tracking to the 3D setting, thus providing an alternative for data association for 3D MOT. Our method outperforms the baseline using one-stage bipartite matching for data association by achieving 0.587 Average Multi-Object Tracking Accuracy (AMOTA) in NuScenes validation set and 0.365 AMOTA (at level 2) in Waymo test set.<\/jats:p>","DOI":"10.3390\/s21092894","type":"journal-article","created":{"date-parts":[[2021,4,21]],"date-time":"2021-04-21T21:25:10Z","timestamp":1619040310000},"page":"2894","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["A Two-Stage Data Association Approach for 3D Multi-Object Tracking"],"prefix":"10.3390","volume":"21","author":[{"given":"Minh-Quan","family":"Dao","sequence":"first","affiliation":[{"name":"LS2N, CNRS, \u00c9cole Centrale de Nantes, 1 Rue de la No\u00eb, 44321 Nantes, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1627-2241","authenticated-orcid":false,"given":"Vincent","family":"Fr\u00e9mont","sequence":"additional","affiliation":[{"name":"LS2N, CNRS, \u00c9cole Centrale de Nantes, 1 Rue de la No\u00eb, 44321 Nantes, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,4,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster r-cnn: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016). Ssd: Single shot multibox detector. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Bewley, A., Ge, Z., Ott, L., Ramos, F., and Upcroft, B. (2016, January 25\u201328). Simple online and realtime tracking. Proceedings of the 2016 IEEE International Conference on Image Processing (ICIP), Phoenix, AZ, USA.","DOI":"10.1109\/ICIP.2016.7533003"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Scheidegger, S., Benjaminsson, J., Rosenberg, E., Krishnan, A., and Granstr\u00f6m, K. (2018, January 26\u201330). Mono-camera 3d multi-object tracking using deep learning detections and pmbm filtering. Proceedings of the 2018 IEEE Intelligent Vehicles Symposium (IV), Changshu, China.","DOI":"10.1109\/IVS.2018.8500454"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Weng, X., Wang, J., Held, D., and Kitani, K. (2020). AB3DMOT: A Baseline for 3D Multi-Object Tracking and New Evaluation Metrics. arXiv.","DOI":"10.1109\/IROS45743.2020.9341164"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"83","DOI":"10.1002\/nav.3800020109","article-title":"The Hungarian method for the assignment problem","volume":"2","author":"Kuhn","year":"1955","journal-title":"Nav. Res. Logist. Q."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Liang, M., Yang, B., Zeng, W., Chen, Y., Hu, R., Casas, S., and Urtasun, R. (2020, January 16\u201318). PnPNet: End-to-End Perception and Prediction with Tracking in the Loop. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01157"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Yin, T., Zhou, X., and Kr\u00e4henb\u00fchl, P. (2020). Center-based 3d object detection and tracking. arXiv.","DOI":"10.1109\/CVPR46437.2021.01161"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Luo, W., Yang, B., and Urtasun, R. (2018, January 18\u201323). Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net. Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00376"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The kitti dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. (2020, January 13\u201319). Nuscenes: A multimodal dataset for autonomous driving. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., and Caine, B. (2020, January 14\u201319). Scalability in perception for autonomous driving: Waymo open dataset. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00252"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1883","DOI":"10.1109\/TAES.2018.2805153","article-title":"Poisson Multi-Bernoulli Mixture Filter: Direct Derivation and Implementation","volume":"54","author":"Williams","year":"2018","journal-title":"IEEE Trans. Aerosp. Electron. Syst."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"kuang Chiu, H., Prioletti, A., Li, J., and Bohg, J. (2020). Probabilistic 3D Multi-Object Tracking for Autonomous Driving. arXiv.","DOI":"10.1109\/ICRA48506.2021.9561754"},{"key":"ref_16","unstructured":"Zhu, B., Jiang, Z., Zhou, X., Li, Z., and Yu, G. (2019). Class-balanced grouping and sampling for point cloud 3d object detection. arXiv."},{"key":"ref_17","unstructured":"Zhou, X., Wang, D., and Kr\u00e4henb\u00fchl, P. (2019). Objects as Points. arXiv."},{"key":"ref_18","unstructured":"Ding, Z., Hu, Y., Ge, R., Huang, L., Chen, S., Wang, Y., and Liao, J. (2020). 1st Place Solution for Waymo Open Dataset Challenge\u20143D Detection and Domain Adaptation. arXiv."},{"key":"ref_19","unstructured":"Ge, R., Ding, Z., Hu, Y., Wang, Y., Chen, S., Huang, L., and Li, Y. (2020). AFDet: Anchor Free One Stage 3D Object Detection. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Shi, S., Guo, C., Jiang, L., Wang, Z., Shi, J., Wang, X., and Li, H. (2020, January 16\u201318). Pv-rcnn: Point-voxel feature set abstraction for 3d object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01054"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Cheng, S., Leng, Z., Cubuk, E.D., Zoph, B., Bai, C., Ngiam, J., Song, Y., Caine, B., Vasudevan, V., and Li, C. (2020). Improving 3D Object Detection through Progressive Population Based Augmentation. arXiv.","DOI":"10.1007\/978-3-030-58589-1_17"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Bae, S.H., and Yoon, K.J. (2014, January 23\u201328). Robust online multi-object tracking based on tracklet confidence and online discriminative appearance learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.159"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"3782","DOI":"10.1109\/TITS.2019.2892405","article-title":"A survey on 3d object detection methods for autonomous driving applications","volume":"20","author":"Arnold","year":"2019","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1012","DOI":"10.1109\/TPAMI.2013.185","article-title":"3d traffic scene understanding from movable platforms","volume":"36","author":"Geiger","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_25","unstructured":"Leal-Taix\u00e9, L., Milan, A., Reid, I., Roth, S., and Schindler, K. (2015). Motchallenge 2015: Towards a benchmark for multi-target tracking. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Mauri, A., Khemmar, R., Decoux, B., Ragot, N., Rossi, R., Trabelsi, R., Boutteau, R., Ertaud, J.Y., and Savatier, X. (2020). Deep Learning for Real-Time 3D Multi-Object Detection, Localisation, and Tracking: Application to Smart Mobility. Sensors, 20.","DOI":"10.3390\/s20020532"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"595","DOI":"10.1109\/TPAMI.2017.2691769","article-title":"Confidence-based data association and discriminative deep appearance learning for robust online multi-object tracking","volume":"40","author":"Bae","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"4178","DOI":"10.1109\/TII.2019.2897128","article-title":"An efficient edge artificial intelligence multipedestrian tracking method with rank constraint","volume":"15","author":"Yang","year":"2019","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1155\/2008\/246309","article-title":"Evaluating multiple object tracking performance: The CLEAR MOT metrics","volume":"2008","author":"Bernardin","year":"2008","journal-title":"EURASIP J. Image Video Process."},{"key":"ref_30","unstructured":"Leal-Taix\u00e9, L., Milan, A., Schindler, K., Cremers, D., Reid, I., and Roth, S. (2017). Tracking the trackers: An analysis of the state of the art in multiple object tracking. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Leal-Taix\u00e9, L., Canton-Ferrer, C., and Schindler, K. (2016, January 12). Learning by tracking: Siamese CNN for robust target association. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Paris, France.","DOI":"10.1109\/CVPRW.2016.59"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Chang, S., Li, W., Zhang, Y., and Feng, Z. (2019). Online siamese network for visual object tracking. Sensors, 19.","DOI":"10.3390\/s19081858"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Weng, X., Wang, Y., Man, Y., and Kitani, K.M. (2020, January 14\u201319). Gnn3dmot: Graph neural network for 3d multi-object tracking with 2d-3d multi-feature learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00653"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/9\/2894\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:50:32Z","timestamp":1760161832000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/9\/2894"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,21]]},"references-count":33,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2021,5]]}},"alternative-id":["s21092894"],"URL":"https:\/\/doi.org\/10.3390\/s21092894","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,4,21]]}}}