{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,29]],"date-time":"2025-10-29T03:50:42Z","timestamp":1761709842104,"version":"build-2065373602"},"reference-count":53,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2021,3,17]],"date-time":"2021-03-17T00:00:00Z","timestamp":1615939200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Autonomous systems need to localize and track surrounding objects in 3D space for safe motion planning. As a result, 3D multi-object tracking (MOT) plays a vital role in autonomous navigation. Most MOT methods use a tracking-by-detection pipeline, which includes both the object detection and data association tasks. However, many approaches detect objects in 2D RGB sequences for tracking, which lacks reliability when localizing objects in 3D space. Furthermore, it is still challenging to learn discriminative features for temporally consistent detection in different frames, and the affinity matrix is typically learned from independent object features without considering the feature interaction between detected objects in the different frames. To settle these problems, we first employ a joint feature extractor to fuse the appearance feature and the motion feature captured from 2D RGB images and 3D point clouds, and then we propose a novel convolutional operation, named RelationConv, to better exploit the correlation between each pair of objects in the adjacent frames and learn a deep affinity matrix for further data association. We finally provide extensive evaluation to reveal that our proposed model achieves state-of-the-art performance on the KITTI tracking benchmark.<\/jats:p>","DOI":"10.3390\/s21062113","type":"journal-article","created":{"date-parts":[[2021,3,17]],"date-time":"2021-03-17T21:43:31Z","timestamp":1616017411000},"page":"2113","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Relation3DMOT: Exploiting Deep Affinity for 3D Multi-Object Tracking from View Aggregation"],"prefix":"10.3390","volume":"21","author":[{"given":"Can","family":"Chen","sequence":"first","affiliation":[{"name":"School of Aerospace, Transport and Manufacturing, Cranfield University, Bedford MK43 0AL, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Luca","family":"Zanotti Fragonara","sequence":"additional","affiliation":[{"name":"School of Aerospace, Transport and Manufacturing, Cranfield University, Bedford MK43 0AL, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3966-7633","authenticated-orcid":false,"given":"Antonios","family":"Tsourdos","sequence":"additional","affiliation":[{"name":"School of Aerospace, Transport and Manufacturing, Cranfield University, Bedford MK43 0AL, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,3,17]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1803","DOI":"10.1109\/LRA.2020.2969183","article-title":"Track to reconstruct and reconstruct to track","volume":"5","author":"Luiten","year":"2020","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_2","unstructured":"Zhang, W., Zhou, H., Sun, S., Wang, Z., Shi, J., and Loy, C.C. (November, January 27). Robust multi-modality multi-object tracking. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_3","unstructured":"Hu, H.N., Cai, Q.Z., Wang, D., Lin, J., Sun, M., Krahenbuhl, P., Darrell, T., and Yu, F. (November, January 27). Joint monocular 3D vehicle detection and tracking. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Weng, X., Wang, J., Held, D., and Kitani, K. (2020). 3D Multi-Object Tracking: A Baseline and New Evaluation Metrics. arXiv.","DOI":"10.1109\/IROS45743.2020.9341164"},{"key":"ref_5","unstructured":"Bergmann, P., Meinhardt, T., and Leal-Taixe, L. (November, January 27). Tracking without bells and whistles. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zhou, X., Koltun, V., and Kr\u00e4henb\u00fchl, P. (2020). Tracking Objects as Points. arXiv.","DOI":"10.1007\/978-3-030-58548-8_28"},{"key":"ref_7","first-page":"104","article-title":"Deep affinity network for multiple object tracking","volume":"43","author":"Sun","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Wang, Z., Zheng, L., Liu, Y., Li, Y., and Wang, S. (2019). Towards real-time multi-object tracking. arXiv.","DOI":"10.1007\/978-3-030-58621-8_7"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Milan, A., Rezatofighi, S.H., Dick, A., Reid, I., and Schindler, K. (2016). Online multi-target tracking using recurrent neural networks. arXiv.","DOI":"10.1609\/aaai.v31i1.11194"},{"key":"ref_10","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). Pointnet: Deep learning on point sets for 3d classification and segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA."},{"key":"ref_11","unstructured":"Qi, C.R., Yi, L., Su, H., and Guibas, L.J. (2017, January 4\u20139). Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Proceedings of the Advances in neural information processing systems, Long Beach, CA, USA."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"35","DOI":"10.1115\/1.3662552","article-title":"A new approach to linear filtering and prediction problems","volume":"82","author":"Kalman","year":"1960","journal-title":"J. Basic Eng."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"83","DOI":"10.1002\/nav.3800020109","article-title":"The Hungarian method for the assignment problem","volume":"2","author":"Kuhn","year":"1955","journal-title":"Nav. Res. Logist. Q."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Weng, X., Wang, Y., Man, Y., and Kitani, K. (2020). GNN3DMOT: Graph Neural Network for 3D Multi-Object Tracking with Multi-Feature Learning. arXiv.","DOI":"10.1109\/CVPR42600.2020.00653"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Shenoi, A., Patel, M., Gwak, J., Goebel, P., Sadeghian, A., Rezatofighi, H., Martin-Martin, R., and Savarese, S. (2020). JRMOT: A Real-Time 3D Multi-Object Tracker and a New Large-Scale Dataset. arXiv.","DOI":"10.1109\/IROS45743.2020.9341635"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"173","DOI":"10.1109\/JOE.1983.1145560","article-title":"Sonar tracking of multiple targets using joint probabilistic data association","volume":"8","author":"Fortmann","year":"1983","journal-title":"IEEE J. Ocean. Eng."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Ma, C., Huang, J.B., Yang, X., and Yang, M.H. (2015, January 7\u201313). Hierarchical convolutional features for visual tracking. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.352"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Wang, L., Ouyang, W., Wang, X., and Lu, H. (2015, January 7\u201313). Visual tracking with fully convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.357"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Bhat, G., Johnander, J., Danelljan, M., Shahbaz Khan, F., and Felsberg, M. (2018, January 8\u201314). Unveiling the power of deep tracking. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01216-8_30"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Danelljan, M., Bhat, G., Shahbaz Khan, F., and Felsberg, M. (2017, January 21\u201326). Eco: Efficient convolution operators for tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.733"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Danelljan, M., Robinson, A., Khan, F.S., and Felsberg, M. (2016, January 8\u201316). Beyond correlation filters: Learning continuous convolution operators for visual tracking. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46454-1_29"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Choi, J., Jin Chang, H., Yun, S., Fischer, T., Demiris, Y., and Young Choi, J. (2017, January 21\u201326). Attentional correlation filter network for adaptive visual tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.513"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Kiani Galoogahi, H., Fagg, A., and Lucey, S. (2017, January 22\u201329). Learning background-aware correlation filters for visual tracking. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.129"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Mueller, M., Smith, N., and Ghanem, B. (2017, January 21\u201326). Context-aware correlation filter tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.152"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Yun, S., Choi, J., Yoo, Y., Yun, K., and Young Choi, J. (2017, January 21\u201326). Action-decision networks for visual tracking with deep reinforcement learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.148"},{"key":"ref_26","unstructured":"Zhang, L., Li, Y., and Nevatia, R. (2008, January 23\u201328). Global data association for multi-object tracking using network flows. Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition, Anchorage, AK, USA."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Schulter, S., Vernaza, P., Choi, W., and Chandraker, M. (2017, January 21\u201326). Deep network flow for multi-object tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.292"},{"key":"ref_28","unstructured":"Weng, X., and Kitani, K. (2019). A baseline for 3d multi-object tracking. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Osep, A., Mehner, W., Mathias, M., and Leibe, B. (June, January 29). Combined image-and world-space tracking in traffic scenes. Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore.","DOI":"10.1109\/ICRA.2017.7989230"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Scheidegger, S., Benjaminsson, J., Rosenberg, E., Krishnan, A., and Granstr\u00f6m, K. (2018, January 26\u201330). Mono-camera 3d multi-object tracking using deep learning detections and pmbm filtering. Proceedings of the 2018 IEEE Intelligent Vehicles Symposium (IV), Changshu, China.","DOI":"10.1109\/IVS.2018.8500454"},{"key":"ref_31","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Yin, T., Zhou, X., and Kr\u00e4henb\u00fchl, P. (2020). Center-based 3d object detection and tracking. arXiv.","DOI":"10.1109\/CVPR46437.2021.01161"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Luo, W., Yang, B., and Urtasun, R. (2018, January 18\u201322). Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00376"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Butt, A.A., and Collins, R.T. (2013, January 23\u201328). Multi-target tracking by lagrangian relaxation to min-cost network flow. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.241"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Shi, S., Wang, X., and Li, H. (2019, January 15\u201320). Pointrcnn: 3d object proposal generation and detection from point cloud. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00086"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Lang, A.H., Vora, S., Caesar, H., Zhou, L., Yang, J., and Beijbom, O. (2019, January 15\u201320). Pointpillars: Fast encoders for object detection from point clouds. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01298"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Qi, C.R., Liu, W., Wu, C., Su, H., and Guibas, L.J. (2018, January 18\u201322). Frustum pointnets for 3d object detection from rgb-d data. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00102"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Ren, J., Chen, X., Liu, J., Sun, W., Pang, J., Yan, Q., Tai, Y.W., and Xu, L. (2017, January 21\u201326). Accurate single stage detector using recurrent rolling convolution. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.87"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Bell, S., Lawrence Zitnick, C., Bala, K., and Girshick, R. (2016, January 27\u201330). Inside-outside net: Detecting objects in context with skip pooling and recurrent neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.314"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_41","unstructured":"Li, Y., Bu, R., Sun, M., Wu, W., Di, X., and Chen, B. (2018, January 3\u20138). Pointcnn: Convolution on x-transformed points. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_42","unstructured":"Xu, B., Wang, N., Chen, T., and Li, M. (2015). Empirical evaluation of rectified activations in convolutional network. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The kitti dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_44","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are we ready for autonomous driving. Proceedings of the CVPR, Providence, RI, USA."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1155\/2008\/246309","article-title":"Evaluating multiple object tracking performance: The CLEAR MOT metrics","volume":"2008","author":"Bernardin","year":"2008","journal-title":"EURASIP J. Image Video Process."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Li, Y., Huang, C., and Nevatia, R. (2009, January 20\u201325). Learning to associate: Hybridboosted multi-target tracker for crowded scene. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206735"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"104423","DOI":"10.1109\/ACCESS.2019.2932301","article-title":"Multiple object tracking with attention to appearance, structure, motion and size","volume":"7","author":"Karunasekera","year":"2019","journal-title":"IEEE Access"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Sharma, S., Ansari, J.A., Murthy, J.K., and Krishna, K.M. (2018, January 21\u201325). Beyond pixels: Leveraging geometry and shape cues for online multi-object tracking. Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, QLD, Australia.","DOI":"10.1109\/ICRA.2018.8461018"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"G\u00fcnd\u00fcz, G., and Acarman, T. (2018, January 26\u201330). A lightweight online multiple object vehicle tracking method. Proceedings of the 2018 IEEE Intelligent Vehicles Symposium (IV), Changshu, China.","DOI":"10.1109\/IVS.2018.8500386"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"374","DOI":"10.1109\/TITS.2019.2892413","article-title":"Online multi-object tracking using joint domain information in traffic scenarios","volume":"21","author":"Tian","year":"2019","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Xiang, Y., Alahi, A., and Savarese, S. (2015, January 7\u201313). Learning to track: Online multi-object tracking by decision making. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.534"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Frossard, D., and Urtasun, R. (2018, January 21\u201325). End-to-end learning of multi-sensor 3d tracking by detection. Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, QLD, Australia.","DOI":"10.1109\/ICRA.2018.8462884"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Baser, E., Balasubramanian, V., Bhattacharyya, P., and Czarnecki, K. (2019, January 9\u201312). Fantrack: 3d multi-object tracking with feature association network. Proceedings of the 2019 IEEE Intelligent Vehicles Symposium (IV), Paris, France.","DOI":"10.1109\/IVS.2019.8813779"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/6\/2113\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:37:18Z","timestamp":1760161038000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/6\/2113"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,17]]},"references-count":53,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2021,3]]}},"alternative-id":["s21062113"],"URL":"https:\/\/doi.org\/10.3390\/s21062113","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2021,3,17]]}}}