{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,11]],"date-time":"2025-09-11T19:40:38Z","timestamp":1757619638617,"version":"3.44.0"},"reference-count":48,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T00:00:00Z","timestamp":1753315200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T00:00:00Z","timestamp":1753315200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["42401499"],"award-info":[{"award-number":["42401499"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007129","name":"Shandong Provincial Natural Science Foundation","doi-asserted-by":"crossref","award":["ZR2023QD087"],"award-info":[{"award-number":["ZR2023QD087"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"publisher","award":["2024M764316"],"award-info":[{"award-number":["2024M764316"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J. King Saud Univ. Comput. Inf. Sci."],"published-print":{"date-parts":[[2025,8]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>RGBT (visible-thermal) object tracking holds significant value in complex scenarios such as low-light and hazy environments, enabling robust all-weather tracking by leveraging the complementary strengths of visible and thermal infrared modalities. However, challenges such as target appearance variations, similar object interference, and camera motion often lead to tracking drift. This paper proposes RecheckTrack, a robust RGBT tracking framework that addresses these issues through the enhancement of temporal information and a backward trajectory verification mechanism. The dual-branch fusion network adaptively learns target dynamics using appearance tokens and modality tokens. Modality tokens focus on high-quality features and target-probable regions, while appearance tokens track dynamic changes in target appearance, improving robustness against deformation, occlusion, and scale variations. To mitigate drift caused by sudden target or camera motion, a recheck network is introduced, which employs a two-stage candidate box selection method and jointly matches targets using bidirectional tracking consistency and appearance similarity. Additionally, for long-term tracking scenarios where targets may be lost, the recheck network is improved with a path-consistency-based backward trajectory selection method and an approximate global search strategy, efficiently recovering lost targets. Experiments on the VTUAV, LasHeR, and RGBT234 datasets demonstrate that RecheckTrack significantly reduces tracking drift and improves accuracy, providing an effective solution for RGBT tracking in complex scenarios.<\/jats:p>","DOI":"10.1007\/s44443-025-00144-w","type":"journal-article","created":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T14:20:59Z","timestamp":1753366859000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["A robust RGBT tracking framework with temporal information enhancement and backward trajectory verification"],"prefix":"10.1007","volume":"37","author":[{"given":"Huiwei","family":"Shi","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaodong","family":"Mu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2051-7843","authenticated-orcid":false,"given":"Hao","family":"He","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengliang","family":"Zhong","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peng","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,7,24]]},"reference":[{"key":"144_CR1","unstructured":"Aharon N, Orfaig R, Bobrovsky BZ (2022) Bot-sort: Robust associations multi-pedestrian tracking. arXiv preprint 2206.14651"},{"key":"144_CR2","doi-asserted-by":"crossref","unstructured":"Bodla N, Singh B, Chellappa R, Davis LS (2017) Soft-nms\u2013improving object detection with one line of code. In: Proceedings of international conference on computer vision, pp 5561\u20135569","DOI":"10.1109\/ICCV.2017.593"},{"key":"144_CR3","doi-asserted-by":"crossref","unstructured":"Cai Y, Sui X, Gu G, Chen Q (2023) Learning modality feature fusion via transformer for rgbt-tracking. Infrared Phys Technol 133:104819","DOI":"10.1016\/j.infrared.2023.104819"},{"key":"144_CR4","doi-asserted-by":"crossref","unstructured":"Cai W, Liu Q, Wang Y (2024) Hiptrack: Visual tracking with historical prompts. In: Proceedings of IEEE conference on computer vision and pattern recognition, pp 19258\u201319267","DOI":"10.1109\/CVPR52733.2024.01822"},{"key":"144_CR5","doi-asserted-by":"crossref","unstructured":"Cao Z et\u00a0al (2022) Tctrack: Temporal contexts for aerial tracking. In: Proceedings of IEEE conference on computer vision and pattern recognition, pp 14798\u201314808","DOI":"10.1109\/CVPR52688.2022.01438"},{"key":"144_CR6","doi-asserted-by":"publisher","first-page":"927","DOI":"10.1609\/aaai.v38i2.27852","volume":"38","author":"B Cao","year":"2024","unstructured":"Cao B, Guo J, Zhu P, Hu Q (2024) Bi-directional adapter for multimodal tracking. Proceedings of AAAI conference on artificial intelligence (AAAI) 38:927\u2013935","journal-title":"Proceedings of AAAI conference on artificial intelligence (AAAI)"},{"key":"144_CR7","doi-asserted-by":"crossref","unstructured":"Chen YH, Wang CY, Yang CY (2023) Neighbortrack: Single object tracking by bipartite matching with neighbor tracklets and its applications to sports. In: Proceedings of IEEE conference on computer vision and pattern recognition workshops, pp 5139\u20135148","DOI":"10.1109\/CVPRW59228.2023.00542"},{"key":"144_CR8","doi-asserted-by":"publisher","first-page":"723","DOI":"10.1109\/TPDS.2021.3081254","volume":"33","author":"L Cheng","year":"2022","unstructured":"Cheng L, Wang J, Li Y (2022) Vitrack: Efficient tracking on the edge for commodity video surveillance systems. IEEE Trans Parallel Distrib Syst 33:723\u2013735","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"144_CR9","doi-asserted-by":"crossref","unstructured":"Choi J, Kwon J, Lee KM (2020) Visual tracking by tridentalign and context embedding. In: Proceedings of the conference on asian conference on computer vision, pp 504\u2013520","DOI":"10.1007\/978-3-030-69532-3_31"},{"key":"144_CR10","doi-asserted-by":"crossref","unstructured":"Cui Y et al (2024) Target\u2013distractor memory joint tracking algorithm via credit allocation network. Comput Vis Image Understand 247:104075","DOI":"10.1016\/j.cviu.2024.104075"},{"key":"144_CR11","doi-asserted-by":"crossref","unstructured":"Cui Y, Jiang C, Wang L, Wu G (2021) Mixformer: End-to-end tracking with iterative mixed attention. In: Proceedings of IEEE international conference on computer vision, pp 13608\u201313618","DOI":"10.1109\/CVPR52688.2022.01324"},{"key":"144_CR12","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Houlsby N (2021) An image is worth 16x16 words: Transformers for image recognition at scale. In: Proceedings of international conference on learning representations"},{"key":"144_CR13","doi-asserted-by":"publisher","first-page":"873","DOI":"10.1109\/LSP.2023.3295758","volume":"30","author":"S Fan","year":"2023","unstructured":"Fan S, He C, Wei C, Zheng Y, Chen X (2023) Bayesian dumbbell diffusion model for rgbt object tracking with enriched priors. IEEE Signal Process Lett 30:873\u2013877","journal-title":"IEEE Signal Process Lett"},{"key":"144_CR14","doi-asserted-by":"crossref","unstructured":"Gao Y et\u00a0al (2019) Deep adaptive fusion network for high performance rgbt tracking. In: Proceedings of IEEE international conference on computer vision workshops, pp 1\u20139","DOI":"10.1109\/ICCVW.2019.00017"},{"key":"144_CR15","doi-asserted-by":"crossref","unstructured":"Gao S, Zhou C, Zhang J (2023) Generalized relation modeling for transformer tracking. In: Proceedings of IEEE conference on computer vision and pattern recognition, pp 18686\u201318695","DOI":"10.1109\/CVPR52729.2023.01792"},{"key":"144_CR16","doi-asserted-by":"crossref","unstructured":"Girshick R (2015) Fast r-cnn. In: Proceedings of IEEE international conference on computer vision and pattern recognition, pp 1440\u20131448","DOI":"10.1109\/ICCV.2015.169"},{"key":"144_CR17","first-page":"11037","volume":"36","author":"L Huang","year":"2020","unstructured":"Huang L, Zhao X, Huang K (2020) Globaltrack: A simple and strong baseline for long-term tracking. Proceedings of association for the advancement of artificial intelligence (AAAI) 36:11037\u201311044","journal-title":"Proceedings of association for the advancement of artificial intelligence (AAAI)"},{"key":"144_CR18","unstructured":"Ilya L, Frank H (2019) Decoupled weight and frank hutter. In: Proceedings of international conference on learning representations"},{"key":"144_CR19","unstructured":"Kristan M et\u00a0al (2019) The seventh visual object tracking vot2019 challenge results. In: Proceedings of IEEE international conference on computer vision workshops, pp 1\u201336"},{"key":"144_CR20","doi-asserted-by":"publisher","first-page":"83","DOI":"10.1002\/nav.3800020109","volume":"2","author":"HW Kuhn","year":"1955","unstructured":"Kuhn HW (1955) The hungarian method for the assignment problem. Nav Res Logist Q 2:83\u201397","journal-title":"Nav Res Logist Q"},{"key":"144_CR21","unstructured":"Lee D et\u00a0al (2023) Backtrack: Robust template update via backward tracking of candidate template. arXiv preprint 2308.10604"},{"key":"144_CR22","doi-asserted-by":"publisher","first-page":"392","DOI":"10.1109\/TIP.2021.3130533","volume":"31","author":"C Li","year":"2022","unstructured":"Li C et al (2022) Lasher: A large-scale high-diversity benchmark for rgbt tracking. IEEE Trans Image Process 31:392\u2013404","journal-title":"IEEE Trans Image Process"},{"key":"144_CR23","doi-asserted-by":"publisher","first-page":"37","DOI":"10.11834\/jig.220578","volume":"28","author":"C Li","year":"2023","unstructured":"Li C, Lu A, Liu L, Tang J (2023) Multi-modal visual tracking: a survey. J Image Graph 28:37\u201356","journal-title":"J Image Graph"},{"key":"144_CR24","doi-asserted-by":"crossref","unstructured":"Li C, Lu A, Zheng A, Tu Z, Tang J (2019) Multi-adapter rgbt tracking. In: Proceedings of IEEE internatianl conference on computer vision workshops, pp 2262\u20132270","DOI":"10.1109\/ICCVW.2019.00279"},{"key":"144_CR25","doi-asserted-by":"crossref","unstructured":"Liu L, Li C, Xiao Y, Tang J (2023) Quality-aware rgbt tracking via supervised reliability learning and weighted residual guidance. In: Proceedings of ACM international conference on multimedia, pp 3129\u20133137","DOI":"10.1145\/3581783.3612341"},{"key":"144_CR26","doi-asserted-by":"crossref","unstructured":"Long C et\u00a0al (2017) Rgb-t slam: A flexible slam framework by combining appearance and thermal information. In: Proceedings of IEEE international conference on robotics and automation, pp 5682\u20135687","DOI":"10.1109\/ICRA.2017.7989668"},{"key":"144_CR27","doi-asserted-by":"crossref","unstructured":"Luo H, Gu Y, Liao X, Lai S, Jiang W (2019) Bag of tricks and a strong baseline for deep person re-identification. In: Proceedings of IEEE conference on computer vision and pattern recognition workshops, pp 1487\u20131495","DOI":"10.1109\/CVPRW.2019.00190"},{"key":"144_CR28","doi-asserted-by":"crossref","unstructured":"Lu Z, Shuai B, Chen Y, Xu Z, Modolo D (2024) Self-supervised multi-object tracking with path consistency. In: Proceedings of IEEE conference on computer vision and pattern recognition, pp 19016\u201319026","DOI":"10.1109\/CVPR52733.2024.01799"},{"key":"144_CR29","doi-asserted-by":"crossref","unstructured":"Mayer C, Danelljan M, Pani\u00a0Paudel D, Van\u00a0Gool L (2021) Learning target candidate association to keep track of what not to track. In: Proceedings of IEEE conference on computer vision and pattern recognition, pp 13424\u201313434","DOI":"10.1109\/ICCV48922.2021.01319"},{"key":"144_CR30","doi-asserted-by":"crossref","unstructured":"Pe\u00f1a Fern\u00e1ndez CA (2023) An ergodic selection method for kinematic configurations in autonomous, flexible mobile systems. J Intell Robot Syst 109:11","DOI":"10.1007\/s10846-023-01933-z"},{"key":"144_CR31","doi-asserted-by":"crossref","unstructured":"Shi H, Mu X, Shen D, Zhong C (2024) Learning a multimodal feature transformer for rgbt tracking. Signal Image Video Process 18:S239\u2013S250","DOI":"10.1007\/s11760-024-03148-7"},{"key":"144_CR32","doi-asserted-by":"crossref","unstructured":"Voigtlaender P Luiten J, Torr PH, Leibe B (2020) Siam r-cnn: Visual tracking by redetection. In: Proceedings of IEEE conference on computer vision and pattern recognition, pp 6577\u20136587","DOI":"10.1109\/CVPR42600.2020.00661"},{"key":"144_CR33","unstructured":"Wang H et\u00a0al (2024) Temporal adaptive rgbt tracking with modality prompt. arXiv preprint 2401.01244"},{"key":"144_CR34","doi-asserted-by":"crossref","unstructured":"Wang X, Ma W, Zhang K, Li S (2019b) Complexity estimation of infrared image sequence for automatic target track. J Syst Eng Electron 45:3245\u20133255","DOI":"10.1007\/s13369-020-04351-7"},{"key":"144_CR35","doi-asserted-by":"crossref","unstructured":"Wang X, Jabri A, Efros AA (2019a) Learning correspondence from the cycle-consistency of time. In: Proceedings of IEEE conference on computer vision and pattern recognition, pp 2566\u20132576","DOI":"10.1109\/CVPR.2019.00267"},{"key":"144_CR36","doi-asserted-by":"crossref","unstructured":"Yan B, Peng H, Fu J, Wang D, Lu H (2021b) Learning spatio-temporal transformer for visual tracking. In: Proceedings of IEEE international conference on computer vision, pp 10428\u201310437","DOI":"10.1109\/ICCV48922.2021.01028"},{"key":"144_CR37","doi-asserted-by":"crossref","unstructured":"Yan B, Zhang X, Wang D, Lu H, Yang X (2021a) Alpha-refine: Boosting tracking performance by precise bounding box estimation. In: Proceedings of IEEE conference on computer vision and pattern recognition, pp 5289\u20135298","DOI":"10.1109\/CVPR46437.2021.00525"},{"key":"144_CR38","doi-asserted-by":"crossref","unstructured":"Ye B, Chang H, Ma B, Shan S, Chen X (2022) Joint feature learning and relation modeling for tracking: A one-stream framework. In: Proceedings of the conference on european conference on computer vision, pp 341\u2013357","DOI":"10.1007\/978-3-031-20047-2_20"},{"key":"144_CR39","doi-asserted-by":"crossref","unstructured":"Zhang Z et\u00a0al (2021a) Distractor-aware fast tracking via dynamic convolutions and mot philosophy. In: Proceedings of IEEE international conference on computer vision, pp 1024\u20131033","DOI":"10.1109\/CVPR46437.2021.00108"},{"key":"144_CR40","first-page":"1","volume":"13682","author":"Y Zhang","year":"2022","unstructured":"Zhang Y et al (2022) Bytetrack: Multi-object tracking by associating every detection box. Proceedings of IEEE european conference on computer vision 13682:1\u201321","journal-title":"Proceedings of IEEE european conference on computer vision"},{"key":"144_CR41","first-page":"1","volume":"20","author":"Z Zhang","year":"2024","unstructured":"Zhang Z et al (2024) Review and analysis of rgbt single object tracking methods: A fusion perspective. ACM Trans Multimedia Comput Commun Appl 20:1\u201327","journal-title":"ACM Trans Multimedia Comput Commun Appl"},{"key":"144_CR42","doi-asserted-by":"publisher","first-page":"2714","DOI":"10.1007\/s11263-021-01495-3","volume":"129","author":"P Zhang","year":"2021","unstructured":"Zhang P, Wang D, Lu H, Yang X (2021) Learning adaptive attribute-driven representation for real-time rgb-t tracking. Int J Comput Vis 129:2714\u20132729","journal-title":"Int J Comput Vis"},{"key":"144_CR43","doi-asserted-by":"publisher","first-page":"2536","DOI":"10.1007\/s11263-021-01487-3","volume":"129","author":"Y Zhang","year":"2021","unstructured":"Zhang Y, Wang L, Wang D, Qi J, Lu H (2021) Learning regression and verification networks for robust long-term tracking. Int J Comput Vis 129:2536\u20132547","journal-title":"Int J Comput Vis"},{"key":"144_CR44","first-page":"1","volume":"72","author":"F Zhang","year":"2023","unstructured":"Zhang F, Peng H, Yu L, Zhao Y, Chen B (2023) Dual-modality space-time memory network for rgbt tracking. IEEE Trans Instrum Meas 72:1\u201312","journal-title":"IEEE Trans Instrum Meas"},{"key":"144_CR45","doi-asserted-by":"publisher","first-page":"28","DOI":"10.1007\/s00138-024-01508-4","volume":"35","author":"H Zhang","year":"2024","unstructured":"Zhang H, Wang P, Chen Z, Zhang J, Li L (2024) Target\u2013distractor memory joint tracking algorithm via credit allocation network. Mach Vis Appl 35:28","journal-title":"Mach Vis Appl"},{"key":"144_CR46","doi-asserted-by":"crossref","unstructured":"Zhang L, Danelljan M, Gonzalez-Garcia A, Weijer J. vd, Khan FS (2019) Multi-modal fusion for end-to-end rgb-t tracking. In: Proceedings of IEEE international conference on computer vision and pattern recognition, pp 2252\u20132261","DOI":"10.1109\/ICCVW.2019.00278"},{"key":"144_CR47","doi-asserted-by":"crossref","unstructured":"Zhang P, Zhao J, Wang D, Lu H, Ruan X (2022a) Visible-thermal uav tracking: a large-scale benchmark and new baseline. In: Proceedings of IEEE international conference on computer vision, pp 8876\u20138885","DOI":"10.1109\/CVPR52688.2022.00868"},{"key":"144_CR48","doi-asserted-by":"crossref","unstructured":"Zhu Z et\u00a0al (2018) Distractor-aware siamese networks for visual object tracking. In: Proceedings of the conference on european conference on computer vision, pp 103\u2013119","DOI":"10.1007\/978-3-030-01240-3_7"}],"container-title":["Journal of King Saud University Computer and Information Sciences"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44443-025-00144-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44443-025-00144-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44443-025-00144-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,7]],"date-time":"2025-09-07T21:31:15Z","timestamp":1757280675000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44443-025-00144-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,24]]},"references-count":48,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,8]]}},"alternative-id":["144"],"URL":"https:\/\/doi.org\/10.1007\/s44443-025-00144-w","relation":{},"ISSN":["1319-1578","2213-1248"],"issn-type":[{"type":"print","value":"1319-1578"},{"type":"electronic","value":"2213-1248"}],"subject":[],"published":{"date-parts":[[2025,7,24]]},"assertion":[{"value":"22 February 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 June 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 July 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of Interest"}},{"value":"All procedures performed in studies involving human participants were in accordance with the ethical standard and have obtained the informed consent of the relevant people.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical and Informed Consent Statement"}}],"article-number":"132"}}