{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T19:37:52Z","timestamp":1773517072720,"version":"3.50.1"},"reference-count":79,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62302142"],"award-info":[{"award-number":["62302142"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Key Research and Development Programs of China","award":["2023YFC2506800"],"award-info":[{"award-number":["2023YFC2506800"]}]},{"name":"Higher Education Innovation Fund Project","award":["2026A-234"],"award-info":[{"award-number":["2026A-234"]}]},{"name":"National Social Science Fund Project","award":["23XMZ031"],"award-info":[{"award-number":["23XMZ031"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>Infrared Object Tracking (IOT) is challenging due to the low contrast of infrared images, which limits effective spatial feature extraction. Although recent works have explored frequency-domain information, their utilization remains insufficient, and fusion strategies either retain redundancy or fail to fully explore distinctive differences, thus limiting complementary enhancement. To overcome this, we propose a novel tracker that introduces a Target-guided Frequency Transformation Module (TFTM) and a Dual-domain Interactive Fusion Network (DIFN). The former extracts multi-frequency representations across scales and orientations, guided by an adaptive mask strategy to suppress background interference. The latter fuses the two domains with differentiated attention to achieve complementary enhancement. Extensive experiments show that our approach achieves superior performance over state-of-the-art trackers, highlighting the effectiveness of comprehensive frequency-domain integration in IOT.<\/jats:p>","DOI":"10.1145\/3787979","type":"journal-article","created":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T14:07:04Z","timestamp":1768831624000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Infrared Object Tracking via Complementary Dual-domain Interaction with Target-guided Frequency Transformation"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-7632-7490","authenticated-orcid":false,"given":"Pengyu","family":"Huang","sequence":"first","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3818-4277","authenticated-orcid":false,"given":"Jingjing","family":"Wu","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6949-4879","authenticated-orcid":false,"given":"Yanrong","family":"Guo","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5461-3986","authenticated-orcid":false,"given":"Richang","family":"Hong","sequence":"additional","affiliation":[{"name":"Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,3,9]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-48881-3_56"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00628"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","unstructured":"Chun-Fu Chen Quanfu Fan and Rameswar Panda. 2021. CrossViT: Cross-attention multi-scale vision transformer for image classification. arXiv:2103.14899. Retrieved from https:\/\/arxiv.org\/abs\/2103.14899","DOI":"10.1109\/ICCV48922.2021.00041"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTARS.2024.3398361"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01400"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00803"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSII.2022.3223871"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-60639-8_34"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3630100"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.733"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00479"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00721"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1121\/1.406784"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1002\/cpa.3160410705"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01261-8_28"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00687"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2024.3452175","article-title":"Triple-domain feature learning with frequency-aware memory enhancement for moving infrared small target detection","volume":"62","author":"Duan Weiwei","year":"2024","unstructured":"Weiwei Duan, Luping Ji, Shengjia Chen, Sicheng Zhu, and Mao Ye. 2024. Triple-domain feature learning with frequency-aware memory enhancement for moving infrared small target detection. IEEE Transactions on Geoscience and Remote Sensing 62 (2024), 1\u201314. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:270379808","journal-title":"IEEE Transactions on Geoscience and Remote Sensing"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00552"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2024.111665"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01792"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIM.2023.3338701"},{"key":"e_1_3_1_23_2","first-page":"249","volume-title":"Proceedings of the 13th International Conference on Artificial Intelligence and Statistics","author":"Glorot Xavier","year":"2010","unstructured":"Xavier Glorot and Yoshua Bengio. 2010. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics. JMLR Workshop and Conference Proceedings, 249\u2013256."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3140929"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF01456326"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3127357"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2957464"},{"key":"e_1_3_1_29_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Huang Songtao","year":"2025","unstructured":"Songtao Huang, Zhen Zhao, Can Li, and Lei Bai. 2025. TimeKAN: KAN-based frequency decomposition learning architecture for long-term time series forecasting. In Proceedings of the International Conference on Learning Representations (ICLR). Retrieved from https:\/\/openreview.net\/forum?id=wTLc79YNbh"},{"key":"e_1_3_1_30_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision Workshops","author":"Kristan Matej","year":"2015","unstructured":"Matej Kristan, Jiri Matas, Ales Leonardis, Michael Felsberg, Luka Cehovin, Gustavo Fernandez, Tomas Vojir, Gustav Hager, Georg Nebehay, and Roman Pflugfelder. 2015. The visual object tracking vot2015 challenge results. In Proceedings of the IEEE International Conference on Computer Vision Workshops, 1\u201323."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00441"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00935"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2023.102147"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3698399"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2017.07.032"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2018.12.011"},{"key":"e_1_3_1_37_2","unstructured":"Yunfeng Li Bo Wang and Ye Li. 2025. LightFC-X: Lightweight convolutional tracker for RGB-X tracking. arXiv:2502.18143. Retrieved from https:\/\/arxiv.org\/abs\/2502.18143"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2932615"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6828"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6828"},{"key":"e_1_3_1_41_2","first-page":"2114","article-title":"Learning deep multi-level similarity for thermal infrared object tracking","volume":"23","author":"Liu Qiao","year":"2020","unstructured":"Qiao Liu, Xin Li, Zhenyu He, Nana Fan, Di Yuan, and Hongpeng Wang. 2020. Learning deep multi-level similarity for thermal infrared object tracking. IEEE Transactions on Multimedia 23 (2020), 2114\u20132126.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2023.3236895"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2017.07.032"},{"key":"e_1_3_1_44_2","first-page":"1269","article-title":"Learning dual-level deep representation for thermal infrared tracking","volume":"25","author":"Liu Qiao","year":"2022","unstructured":"Qiao Liu, Di Yuan, Nana Fan, Peng Gao, Xin Li, and Zhenyu He. 2022. Learning dual-level deep representation for thermal infrared tracking. IEEE Transactions on Multimedia 25 (2022), 1269\u20131281.","journal-title":"IEEE Transactions on Multimedia"},{"key":"e_1_3_1_45_2","unstructured":"I. Loshchilov. 2017. Decoupled weight decay regularization. arXiv:1711.05101. Retrieved from https:\/\/arxiv.org\/abs\/1711.05101"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00371-022-02765-y"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.3390\/s23156887"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.465"},{"issue":"285","key":"e_1_3_1_49_2","first-page":"23","article-title":"A threshold selection method from gray-level histograms","volume":"11","author":"Otsu Nobuyuki","year":"1975","unstructured":"Nobuyuki Otsu. 1975. A threshold selection method from gray-level histograms. Automatica 11, 285\u2013296 (1975), 23\u201327.","journal-title":"Automatica"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2023.111234"},{"key":"e_1_3_1_51_2","doi-asserted-by":"crossref","unstructured":"Hamid Rezatofighi Nathan Tsoi JunYoung Gwak Amir Sadeghian Ian Reid and Silvio Savarese. 2019. Generalized intersection over union: A metric and a loss for bounding box regression. arXiv:1902.09630. Retrieved from https:\/\/arxiv.org\/abs\/1902.09630","DOI":"10.1109\/CVPR.2019.00075"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00937"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2025.3553695"},{"key":"e_1_3_1_55_2","unstructured":"He Wang Tianyang Xu Zhangyong Tang Xiao-Jun Wu and Josef Kittler. 2025. UASTrack: A unified adaptive selection framework with modality-customization in single object tracking. arXiv:2502.18220. Retrieved from https:\/\/arxiv.org\/abs\/2502.18220"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00142"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00935"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3497746"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2023.3276357"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i8.32938"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00181"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00525"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3074239"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00676"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2022.108485"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2024.3512551"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2024.128908"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2025.107707"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2879249"},{"key":"e_1_3_1_70_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Zhang Qianru","year":"2025","unstructured":"Qianru Zhang, Chenglei Yu, Haixin Wang, Yudong Yan, Yuansheng Cao, Hongzhi Yin, Siu Ming Yiu, and Tailin Wu. 2025. FLDmamba: Integrating Fourier and Laplace transform decomposition with mamba for enhanced time series prediction. In Proceedings of the International Conference on Learning Representations (ICLR). Retrieved from https:\/\/openreview.net\/forum?id=9EiWIyJMNi"},{"key":"e_1_3_1_71_2","doi-asserted-by":"crossref","unstructured":"Shang Zhang Xiaobo Ding Huanbin Zhang Ruoyan Xiong and Yue Zhang. 2025. STARS: Sparse learning correlation filter with spatio-temporal regularization and super-resolution reconstruction for thermal infrared target tracking. arXiv:2504.14491. Retrieved from https:\/\/arxiv.org\/abs\/2504.14491","DOI":"10.1007\/978-981-96-9863-9_6"},{"key":"e_1_3_1_72_2","doi-asserted-by":"crossref","unstructured":"Shang Zhang Huipan Guan Xiaobo Ding Ruoyan Xiong and Yue Zhang. 2025. SMTT: Novel structured multi-task tracking with graph-regularized sparse representation for robust thermal infrared target tracking. arXiv:2504.14566. Retrieved from https:\/\/arxiv.org\/abs\/2504.14566","DOI":"10.1007\/978-981-96-9812-7_28"},{"key":"e_1_3_1_73_2","doi-asserted-by":"crossref","unstructured":"Shang Zhang Huanbin Zhang Dali Feng Yujie Cui Ruoyan Xiong and Cen He. 2025. SMMT: Siamese motion mamba with self-attention for thermal infrared target tracking. arXiv:2505.04088. Retrieved from https:\/\/arxiv.org\/abs\/2505.04088","DOI":"10.1007\/978-981-96-9805-9_20"},{"key":"e_1_3_1_74_2","first-page":"114361","article-title":"Learning diverse fine-grained features for thermal infrared tracking","volume":"168","author":"Zhang Y.","year":"2021","unstructured":"Y. Zhang, L. Chen, and H. Wang. 2021. Learning diverse fine-grained features for thermal infrared tracking. Expert Systems with Applications 168 (2021), 114361.","journal-title":"Expert Systems with Applications"},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00472"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58589-1_46"},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.3390\/rs14010177"},{"key":"e_1_3_1_78_2","first-page":"108","volume-title":"Proceedings of the Chinese Conference on Pattern Recognition and Computer Vision (PRCV)","author":"Zhao Manqi","year":"2023","unstructured":"Manqi Zhao, Shenyang Li, and Han Wang. 2023. Frequency and spatial domain filter network for visual object tracking. In Proceedings of the Chinese Conference on Pattern Recognition and Computer Vision (PRCV). Springer, 108\u2013120."},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3368112"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3050073"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3787979","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T13:24:54Z","timestamp":1773494694000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3787979"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,9]]},"references-count":79,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3787979"],"URL":"https:\/\/doi.org\/10.1145\/3787979","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,9]]},"assertion":[{"value":"2025-09-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-03","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}