{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T17:07:43Z","timestamp":1779210463541,"version":"3.51.4"},"reference-count":43,"publisher":"Walter de Gruyter GmbH","issue":"1","license":[{"start":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T00:00:00Z","timestamp":1767225600000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026,1,23]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Multiple-object tracking (MOT) is a fundamental task in computer vision with many applications. For practical operations, tracking for monitoring with thermal imaging, unaffected by lighting conditions, is important. However, most MOT methods are proposed to analyze video streams from RGB cameras, while there are few datasets and research on multi-object tracking in infrared image sequences. In this paper, we provide a new infrared dataset for object detection and tracking, which contains small objects and occlusion challenges. We also propose a new robust tracker, which enhances object detection with the strategic integration of the convolutional block attention module (CBAM) into the YOLOv7 model, along with specialized fusion of IoU, Size, and ReID features during data association to overcome the challenges of thermal images. Our tracker achieves 59.29 HOTA, 73.46 MOTA, and 74.4 IDF1 as a new state-of-the-art on the CAMEL benchmark. The tracker\u2019s source code and dataset are publicly available at:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/aquarter147\/TMTV_Thermal_MOT\">https:\/\/github.com\/aquarter147\/TMTV_Thermal_MOT<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1515\/comp-2025-0055","type":"journal-article","created":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T16:23:56Z","timestamp":1779207836000},"source":"Crossref","is-referenced-by-count":0,"title":["Advancing thermal multi-object tracking with attention and\u00a0metric fusion"],"prefix":"10.1515","volume":"16","author":[{"given":"Thao-Anh","family":"Tran","sequence":"first","affiliation":[{"name":"School of Electrical and Electronic Engineering , Hanoi University of Science and Technology , Hanoi , Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vu-Minh","family":"Le","sequence":"additional","affiliation":[{"name":"Optoelectronics Center , Viettel Aerospace Institute, Viettel Group , Hanoi , Vietnam"},{"name":"University of Engineering and Technology, Vietnam National University , Hanoi , Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Thanh-Tung","family":"Phan","sequence":"additional","affiliation":[{"name":"Optoelectronics Center , Viettel Aerospace Institute, Viettel Group , Hanoi , Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dung","family":"Hoang","sequence":"additional","affiliation":[{"name":"Optoelectronics Center , Viettel Aerospace Institute, Viettel Group , Hanoi , Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Duc","family":"Phan","sequence":"additional","affiliation":[{"name":"University of Engineering and Technology, Vietnam National University , Hanoi , Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huong","family":"Ninh","sequence":"additional","affiliation":[{"name":"Optoelectronics Center , Viettel Aerospace Institute, Viettel Group , Hanoi , Vietnam"},{"name":"University of Engineering and Technology, Vietnam National University , Hanoi , Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hai","family":"Tran","sequence":"additional","affiliation":[{"name":"Optoelectronics Center , Viettel Aerospace Institute, Viettel Group , Hanoi , Vietnam"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"374","published-online":{"date-parts":[[2026,5,19]]},"reference":[{"key":"2026051916235012631_j_comp-2025-0055_ref_001","doi-asserted-by":"crossref","unstructured":"I. Shopovska, L. Jovanov, and W. Philips, \u201cDeep visible and thermal image fusion for enhanced pedestrian visibility,\u201d Sensors, vol.\u00a019, no.\u00a017, p.\u00a03727, 2019, https:\/\/doi.org\/10.3390\/s19173727.","DOI":"10.3390\/s19173727"},{"key":"2026051916235012631_j_comp-2025-0055_ref_002","doi-asserted-by":"crossref","unstructured":"E. Tosun, O. F. Dinc, B. Arli, and S. Tozburun, \u201cA classifier for dynamic thermal imaging,\u201d in European Conference on Biomedical Optics, Optica Publishing Group, 2023, p.\u00a0126271H.","DOI":"10.1117\/12.2672244"},{"key":"2026051916235012631_j_comp-2025-0055_ref_003","doi-asserted-by":"crossref","unstructured":"N. Wojke, A. Bewley, and D. Paulus, \u201cSimple online and realtime tracking with a deep association metric,\u201d in 2017 IEEE International Conference on Image Processing (ICIP), IEEE, 2017, pp.\u00a03645\u20133649.","DOI":"10.1109\/ICIP.2017.8296962"},{"key":"2026051916235012631_j_comp-2025-0055_ref_004","doi-asserted-by":"crossref","unstructured":"Y. Zhang, C. Wang, X. Wang, W. Zeng, and W. Liu, \u201cFairmot: On the fairness of detection and re-identification in multiple object tracking,\u201d Int. J. Comput. Vis., vol.\u00a0129, no.\u00a011, pp.\u00a03069\u20133087, 2021, https:\/\/doi.org\/10.1007\/s11263-021-01513-4.","DOI":"10.1007\/s11263-021-01513-4"},{"key":"2026051916235012631_j_comp-2025-0055_ref_005","doi-asserted-by":"crossref","unstructured":"Z. Lu, V. Rathod, R. Votel, and J. Huang, \u201cRetinatrack: Online single stage joint detection and tracking,\u201d in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp.\u00a014668\u201314678.","DOI":"10.1109\/CVPR42600.2020.01468"},{"key":"2026051916235012631_j_comp-2025-0055_ref_006","doi-asserted-by":"crossref","unstructured":"C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, \u201cYolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,\u201d in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp.\u00a07464\u20137475.","DOI":"10.1109\/CVPR52729.2023.00721"},{"key":"2026051916235012631_j_comp-2025-0055_ref_007","doi-asserted-by":"crossref","unstructured":"S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, \u201cCbam: Convolutional block attention module,\u201d in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp.\u00a03\u201319.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"2026051916235012631_j_comp-2025-0055_ref_008","doi-asserted-by":"crossref","unstructured":"E. Gebhardt and M. Wolf, \u201cCamel dataset for visual and thermal infrared multiple object detection and tracking,\u201d in 2018 15th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS), IEEE, 2018, pp.\u00a01\u20136.","DOI":"10.1109\/AVSS.2018.8639094"},{"key":"2026051916235012631_j_comp-2025-0055_ref_009","doi-asserted-by":"crossref","unstructured":"W. A. El Ahmar, D. Kolhatkar, F. E. Nowruzi, H. AlGhamdi, J. Hou, and R. Laganiere, \u201cMultiple object detection and tracking in the thermal spectrum,\u201d in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp.\u00a0277\u2013285.","DOI":"10.1109\/CVPRW56347.2022.00042"},{"key":"2026051916235012631_j_comp-2025-0055_ref_010","doi-asserted-by":"crossref","unstructured":"R. Girshick, J. Donahue, T. Darrell, and J. Malik, \u201cRich feature hierarchies for accurate object detection and semantic segmentation,\u201d in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp.\u00a0580\u2013587.","DOI":"10.1109\/CVPR.2014.81"},{"key":"2026051916235012631_j_comp-2025-0055_ref_011","doi-asserted-by":"crossref","unstructured":"R. Girshick, \u201cFast r-cnn,\u201d in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp.\u00a01440\u20131448.","DOI":"10.1109\/ICCV.2015.169"},{"key":"2026051916235012631_j_comp-2025-0055_ref_012","doi-asserted-by":"crossref","unstructured":"S. Ren, K. He, R. Girshick, and J. Sun, \u201cFaster r-cnn: Towards real-time object detection with region proposal networks,\u201d IEEE Trans. Pattern Anal. Mach. Intell., vol.\u00a039, no.\u00a06, pp.\u00a01137\u20131149, 2016, https:\/\/doi.org\/10.1109\/tpami.2016.2577031.","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"2026051916235012631_j_comp-2025-0055_ref_013","doi-asserted-by":"crossref","unstructured":"K. He, G. Gkioxari, P. Doll\u00e1r, and R. Girshick, \u201cMask r-cnn,\u201d in Proceedings of the IEEE International Conference on Computer Vision, 2017, pp.\u00a02961\u20132969.","DOI":"10.1109\/ICCV.2017.322"},{"key":"2026051916235012631_j_comp-2025-0055_ref_014","doi-asserted-by":"crossref","unstructured":"K. He, X. Zhang, S. Ren, and J. Sun, \u201cDeep residual learning for image recognition,\u201d in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp.\u00a0770\u2013778.","DOI":"10.1109\/CVPR.2016.90"},{"key":"2026051916235012631_j_comp-2025-0055_ref_015","unstructured":"K. Simonyan and A. Zisserman, \u201cVery deep convolutional networks for large-scale image recognition,\u201d 2015, https:\/\/arxiv.org\/abs\/1409.1556."},{"key":"2026051916235012631_j_comp-2025-0055_ref_016","doi-asserted-by":"crossref","unstructured":"J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, \u201cYou only look once: Unified, real-time object detection,\u201d in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp.\u00a0779\u2013788.","DOI":"10.1109\/CVPR.2016.91"},{"key":"2026051916235012631_j_comp-2025-0055_ref_017","doi-asserted-by":"crossref","unstructured":"N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, \u201cEnd-to-end object detection with transformers,\u201d in European Conference on Computer Vision, Springer, 2020, pp.\u00a0213\u2013229.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"2026051916235012631_j_comp-2025-0055_ref_018","unstructured":"H. Zhang et al.., \u201cDino: Detr with improved denoising anchor boxes for end-to-end object detection,\u201d arXiv preprint arXiv:2203.03605, 2022."},{"key":"2026051916235012631_j_comp-2025-0055_ref_019","doi-asserted-by":"crossref","unstructured":"Y. Li, S. Li, H. Du, L. Chen, D. Zhang, and Y. Li, \u201cYolo-acn: Focusing on small target and occluded object detection,\u201d IEEE Access, vol.\u00a08, pp.\u00a0227 288\u2013227 303, 2020, https:\/\/doi.org\/10.1109\/access.2020.3046515.","DOI":"10.1109\/ACCESS.2020.3046515"},{"key":"2026051916235012631_j_comp-2025-0055_ref_020","doi-asserted-by":"crossref","unstructured":"A. Bewley, Z. Ge, L. Ott, F. Ramos, and B. Upcroft, \u201cSimple online and realtime tracking,\u201d in 2016 IEEE International Conference on Image Processing (ICIP), IEEE, 2016, pp.\u00a03464\u20133468.","DOI":"10.1109\/ICIP.2016.7533003"},{"key":"2026051916235012631_j_comp-2025-0055_ref_021","doi-asserted-by":"crossref","unstructured":"H. W. Kuhn, \u201cThe Hungarian method for the assignment problem,\u201d Nav. Res. Logist. Q., vol.\u00a02, nos. 1\u20132, pp.\u00a083\u201397, 1955, https:\/\/doi.org\/10.1002\/nav.3800020109.","DOI":"10.1002\/nav.3800020109"},{"key":"2026051916235012631_j_comp-2025-0055_ref_022","doi-asserted-by":"crossref","unstructured":"Y. Zhang et al.., \u201cBytetrack: Multi-object tracking by associating every detection box,\u201d in European Conference on Computer Vision, Springer, 2022, pp.\u00a01\u201321.","DOI":"10.1007\/978-3-031-20047-2_1"},{"key":"2026051916235012631_j_comp-2025-0055_ref_023","unstructured":"N. Aharon, R. Orfaig, and B.-Z. Bobrovsky, \u201cBot-sort: Robust associations multi-pedestrian tracking,\u201d arXiv preprint arXiv:2206.14651, 2022."},{"key":"2026051916235012631_j_comp-2025-0055_ref_024","unstructured":"A. Milan, \u201cMot16: A benchmark for multi-object tracking,\u201d arXiv preprint arXiv:1603.00831, 2016."},{"key":"2026051916235012631_j_comp-2025-0055_ref_025","unstructured":"P. Dendorfer, \u201cMot20: A benchmark for multi object tracking in crowded scenes,\u201d arXiv preprint arXiv:2003.09003, 2020."},{"key":"2026051916235012631_j_comp-2025-0055_ref_026","doi-asserted-by":"crossref","unstructured":"F. Yu et al.., \u201cBdd100k: A diverse driving dataset for heterogeneous multitask learning,\u201d in Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp.\u00a02636\u20132645.","DOI":"10.1109\/CVPR42600.2020.00271"},{"key":"2026051916235012631_j_comp-2025-0055_ref_027","unstructured":"Teledyne, FLIR. Free Teledyne Flir Thermal Dataset for Algorithm Training, 2023. Available at: https:\/\/www.flir.com\/oem\/adas\/adas-dataset-form\/ [Accessed: Nov. 13, 2023]."},{"key":"2026051916235012631_j_comp-2025-0055_ref_028","doi-asserted-by":"crossref","unstructured":"S. Du, B. Zhang, P. Zhang, P. Xiang, and H. Xue, \u201cFa-yolo: An improved yolo model for infrared occlusion object detection under confusing background,\u201d Wireless Commun. Mobile Comput., vol.\u00a02021, no.\u00a01, p.\u00a01896029, 2021, https:\/\/doi.org\/10.1155\/2021\/1896029.","DOI":"10.1155\/2021\/1896029"},{"key":"2026051916235012631_j_comp-2025-0055_ref_029","doi-asserted-by":"crossref","unstructured":"M. Sun, H. Zhang, Z. Huang, Y. Luo, and Y. Li, \u201cRoad infrared target detection with i-yolo,\u201d IET Image Process., vol.\u00a016, no.\u00a01, pp.\u00a092\u2013101, 2022, https:\/\/doi.org\/10.1049\/ipr2.12331.","DOI":"10.1049\/ipr2.12331"},{"key":"2026051916235012631_j_comp-2025-0055_ref_030","doi-asserted-by":"crossref","unstructured":"X. Dai, X. Yuan, and X. Wei, \u201cTirnet: Object detection in thermal infrared images for autonomous driving,\u201d Appl. Intell., vol.\u00a051, no.\u00a03, pp.\u00a01244\u20131261, 2021, https:\/\/doi.org\/10.1007\/s10489-020-01882-2.","DOI":"10.1007\/s10489-020-01882-2"},{"key":"2026051916235012631_j_comp-2025-0055_ref_031","doi-asserted-by":"crossref","unstructured":"M. P. Muresan, S. Nedevschi, and R. Danescu, \u201cRobust data association using fusion of data-driven and engineered features for real-time pedestrian tracking in thermal images,\u201d Sensors, vol.\u00a021, no.\u00a023, p.\u00a08005, 2021, https:\/\/doi.org\/10.3390\/s21238005.","DOI":"10.3390\/s21238005"},{"key":"2026051916235012631_j_comp-2025-0055_ref_032","doi-asserted-by":"crossref","unstructured":"M. P. Muresan, R. Danescu, and S. Nedevschi, \u201cMulti-object tracking, segmentation and validation in thermal images,\u201d in 2023 IEEE Intelligent Vehicles Symposium (IV), IEEE, 2023, pp.\u00a01\u20138.","DOI":"10.1109\/IV55152.2023.10186655"},{"key":"2026051916235012631_j_comp-2025-0055_ref_033","doi-asserted-by":"crossref","unstructured":"Y. Li, P. Wei, M. You, Y. Wei, and H. Zhang, \u201cJoint detection, tracking, and classification of multiple extended objects based on the jdtc-pmbm-ggiw filter,\u201d Remote Sens., vol.\u00a015, no.\u00a04, p.\u00a0887, 2023, https:\/\/doi.org\/10.3390\/rs15040887.","DOI":"10.3390\/rs15040887"},{"key":"2026051916235012631_j_comp-2025-0055_ref_034","doi-asserted-by":"crossref","unstructured":"N. Ibrahim, A. R. Darlis, and B. Kusumoputro, \u201cPerformance analysis of yolo-deep sort on thermal video-based online multi-objet tracking,\u201d in 2023 IEEE 13th International Conference on Consumer Electronics-Berlin (ICCE-Berlin), IEEE, 2023, pp.\u00a01\u20136.","DOI":"10.1109\/ICCE-Berlin58801.2023.10375683"},{"key":"2026051916235012631_j_comp-2025-0055_ref_035","doi-asserted-by":"crossref","unstructured":"N. Ibrahim, A. Ramadhan Darlis, A. Subiantoro, F. Yusivar, and B. Kusumoputro, \u201cOnline multi-object tracking (mot) of sequential thermal images based on deep appearance features using yolo-deepsort,\u201d SSRN Electron. J., 2023, https:\/\/doi.org\/10.2139\/ssrn.4364548.","DOI":"10.2139\/ssrn.4364548"},{"key":"2026051916235012631_j_comp-2025-0055_ref_036","unstructured":"C.-Y. Wang, H.-Y. M. Liao, and I.-H. Yeh, \u201cDesigning network design strategies through gradient path analysis,\u201d 2022, https:\/\/arxiv.org\/abs\/2211.04800."},{"key":"2026051916235012631_j_comp-2025-0055_ref_037","doi-asserted-by":"crossref","unstructured":"H. Luo, Y. Gu, X. Liao, S. Lai, and W. Jiang, \u201cBag of tricks and a strong baseline for deep person re-identification,\u201d 2019, https:\/\/arxiv.org\/abs\/1903.07071.","DOI":"10.1109\/CVPRW.2019.00190"},{"key":"2026051916235012631_j_comp-2025-0055_ref_038","doi-asserted-by":"crossref","unstructured":"L. He, X. Liao, W. Liu, X. Liu, P. Cheng, and T. Mei, \u201cFastreid: A pytorch toolbox for general instance re-identification,\u201d in Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp.\u00a09664\u20139667.","DOI":"10.1145\/3581783.3613460"},{"key":"2026051916235012631_j_comp-2025-0055_ref_039","doi-asserted-by":"crossref","unstructured":"Z. Wang, L. Zheng, Y. Liu, Y. Li, and S. Wang, \u201cTowards real-time multi-object tracking,\u201d in European Conference on Computer Vision, Springer, 2020, pp.\u00a0107\u2013122.","DOI":"10.1007\/978-3-030-58621-8_7"},{"key":"2026051916235012631_j_comp-2025-0055_ref_040","unstructured":"M. Zhu, \u201cRecall, precision and average precision,\u201d Dept. Statist. Actuarial Sci., Univ. Waterloo, Waterloo, ON, Canada, Tech. Rep. 2004\u201309, 2004."},{"key":"2026051916235012631_j_comp-2025-0055_ref_041","doi-asserted-by":"crossref","unstructured":"K. Bernardin and R. Stiefelhagen, \u201cEvaluating multiple object tracking performance: The clear mot metrics,\u201d EURASIP J. Image Video Process., vol. 2008, no. 1, pp. 1\u201310, 2008. https:\/\/doi.org\/10.1155\/2008\/246309.","DOI":"10.1155\/2008\/246309"},{"key":"2026051916235012631_j_comp-2025-0055_ref_042","doi-asserted-by":"crossref","unstructured":"E. Ristani, F. Solera, R. Zou, R. Cucchiara, and C. Tomasi, \u201cPerformance measures and a data set for multi-target, multi-camera tracking,\u201d in European Conference on Computer Vision, Springer, 2016, pp.\u00a017\u201335.","DOI":"10.1007\/978-3-319-48881-3_2"},{"key":"2026051916235012631_j_comp-2025-0055_ref_043","doi-asserted-by":"crossref","unstructured":"J. Luiten et al.., \u201cHota: A higher order metric for evaluating multi-object tracking,\u201d Int. J. Comput. Vis., vol. 129, no. 2, pp. 548\u2013578, 2021. https:\/\/doi.org\/10.1007\/s11263-020-01375-2.","DOI":"10.1007\/s11263-020-01375-2"}],"container-title":["Open Computer Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/comp-2025-0055\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/comp-2025-0055\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T16:24:06Z","timestamp":1779207846000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/comp-2025-0055\/html"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,1]]},"references-count":43,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,5,19]]},"published-print":{"date-parts":[[2026,1,23]]}},"alternative-id":["10.1515\/comp-2025-0055"],"URL":"https:\/\/doi.org\/10.1515\/comp-2025-0055","relation":{},"ISSN":["2299-1093"],"issn-type":[{"value":"2299-1093","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,1]]},"article-number":"20250055"}}