{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,2]],"date-time":"2025-08-02T16:49:08Z","timestamp":1754153348380,"version":"3.41.2"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61906168, 62201400, and 62272267"],"award-info":[{"award-number":["61906168, 62201400, and 62272267"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100022955","name":"Fundamental Research Funds for the Provincial Universities of Zhejiang","doi-asserted-by":"crossref","award":["RF-A2024013"],"award-info":[{"award-number":["RF-A2024013"]}],"id":[{"id":"10.13039\/100022955","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Zhejiang Provincial Natural Science Foundation of China","award":["LY23F020023, LZ23F020001"],"award-info":[{"award-number":["LY23F020023, LZ23F020001"]}]},{"name":"Construction of Hubei Provincial Key Laboratory for Intelligent Visual Monitoring of Hydropower Projects","award":["2022SDSJ01"],"award-info":[{"award-number":["2022SDSJ01"]}]},{"name":"Hangzhou AI major scientific and technological innovation project","award":["2022AIZD0061"],"award-info":[{"award-number":["2022AIZD0061"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2025,8,31]]},"abstract":"<jats:p>\n            The current one-stream tracking pipelines are early relation modeling in feature extraction. However, insufficient discrimination may result in ambiguous relation modeling during early feature extraction. Moreover, the non-target information occupies most of the search image, rendering most relation modeling futile. To tackle the above issues, we propose tracking via learning adaptive target-oriented representation, named\n            <jats:italic toggle=\"yes\">ATOTrack<\/jats:italic>\n            . We design an\n            <jats:italic toggle=\"yes\">Untied positional encoding<\/jats:italic>\n            to mark the template token and the search region token separately, which reduces the confused relationship between the template and the search region. Besides, we introduce an Auto-Mask Learner to decouple the target and non-target information in the search region. Interestingly, the Auto-Mask Learner can self-learn and mask the ineffective information to interpret adaptive target-oriented representation. Extensive experiments demonstrate that ATOTrack is superior to existing methods, which achieves the state-of-the-art performance on six tracking benchmarks. In particular, ATOTrack establishes a new record on AViST with\n            <jats:italic toggle=\"yes\">57%<\/jats:italic>\n            AO. The code and models will be released as soon.\n          <\/jats:p>","DOI":"10.1145\/3732785","type":"journal-article","created":{"date-parts":[[2025,4,29]],"date-time":"2025-04-29T12:54:44Z","timestamp":1745931284000},"page":"1-15","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Adaptive Target-Oriented Tracking"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8916-1174","authenticated-orcid":false,"given":"Sixian","family":"Chan","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology, Zhejiang University of Technology, Hangzhou, China and Hubei Key Laboratory of Intelligent Vision Based Monitoring for Hydroelectric Engineering, the College of Computer and Information, China Three Gorges University, Yichang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-0529-3593","authenticated-orcid":false,"given":"Xianpeng","family":"Zeng","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Zhejiang University of Technology, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-2452-8655","authenticated-orcid":false,"given":"Zhoujian","family":"Wu","sequence":"additional","affiliation":[{"name":"Hangzhou Xsuan Technology Co., Ltd, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-6125-5933","authenticated-orcid":false,"given":"Yu","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Zhejiang University of Technology, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0732-5169","authenticated-orcid":false,"given":"Xiaolong","family":"Zhou","sequence":"additional","affiliation":[{"name":"The College of Electrical and Information Engineering, Quzhou University, Quzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6301-3592","authenticated-orcid":false,"given":"Tinglong","family":"Tang","sequence":"additional","affiliation":[{"name":"Hubei Key Laboratory of Intelligent Vision Based Monitoring for Hydroelectric Engineering, The College of Computer and Information, China Three Gorges University, Yichang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3296-5459","authenticated-orcid":false,"given":"Jie","family":"Hu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Intelligent Informatics for Safety &amp; Emergency of Zhejiang Province, Wenzhou University, Wenzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,7,22]]},"reference":[{"doi-asserted-by":"publisher","key":"e_1_3_1_2_2","DOI":"10.1145\/1577769.1577771"},{"doi-asserted-by":"publisher","key":"e_1_3_1_3_2","DOI":"10.1007\/978-3-319-48881-3_56"},{"doi-asserted-by":"publisher","key":"e_1_3_1_4_2","DOI":"10.1109\/ICCV.2019.00628"},{"doi-asserted-by":"publisher","key":"e_1_3_1_5_2","DOI":"10.1109\/WACV56688.2023.00162"},{"doi-asserted-by":"publisher","key":"e_1_3_1_6_2","DOI":"10.1007\/978-3-031-20047-2_37"},{"doi-asserted-by":"publisher","key":"e_1_3_1_7_2","DOI":"10.1109\/ICCV51070.2023.00879"},{"doi-asserted-by":"publisher","key":"e_1_3_1_8_2","DOI":"10.1007\/978-3-031-20047-2_22"},{"doi-asserted-by":"publisher","key":"e_1_3_1_9_2","DOI":"10.1007\/978-3-031-25085-9_26"},{"doi-asserted-by":"publisher","unstructured":"Xin Chen Houwen Peng Dong Wang Huchuan Lu and Han Hu. 2023. SeqTrack: Sequence to sequence learning for visual object tracking. In 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 14572\u201314581. DOI: 10.1109\/CVPR52729.2023.01400","key":"e_1_3_1_10_2","DOI":"10.1109\/CVPR52729.2023.01400"},{"doi-asserted-by":"publisher","key":"e_1_3_1_11_2","DOI":"10.1109\/CVPR46437.2021.00803"},{"unstructured":"Bowen Cheng Anwesa Choudhuri Ishan Misra Alexander Kirillov Rohit Girdhar and Alexander G. Schwing. 2021. Mask2Former for video instance segmentation. arXiv:2112.10764. Retrieved from https:\/\/arxiv.org\/abs\/2112.10764","key":"e_1_3_1_12_2"},{"doi-asserted-by":"publisher","key":"e_1_3_1_13_2","DOI":"10.1109\/CVPR52688.2022.01324"},{"doi-asserted-by":"publisher","key":"e_1_3_1_14_2","DOI":"10.1109\/TPAMI.2024.3349519"},{"doi-asserted-by":"publisher","key":"e_1_3_1_15_2","DOI":"10.1109\/CVPR.2019.00479"},{"unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In North American Chapter of the Association for Computational Linguistics. Retrieved from https:\/\/api.semanticscholar.org\/CorpusID:52967399","key":"e_1_3_1_16_2"},{"key":"e_1_3_1_17_2","volume-title":"9th International Conference on Learning Representations (ICLR \u201921)","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations (ICLR \u201921). OpenReview.net."},{"doi-asserted-by":"publisher","key":"e_1_3_1_18_2","DOI":"10.1007\/s11263-020-01387-y"},{"doi-asserted-by":"publisher","key":"e_1_3_1_19_2","DOI":"10.1109\/CVPR.2019.00552"},{"doi-asserted-by":"publisher","key":"e_1_3_1_20_2","DOI":"10.24963\/ijcai.2022\/127"},{"doi-asserted-by":"publisher","key":"e_1_3_1_21_2","DOI":"10.1007\/978-3-031-20047-2_9"},{"doi-asserted-by":"publisher","key":"e_1_3_1_22_2","DOI":"10.1109\/CVPR52729.2023.01792"},{"doi-asserted-by":"publisher","key":"e_1_3_1_23_2","DOI":"10.1109\/WACV57701.2024.00657"},{"doi-asserted-by":"publisher","key":"e_1_3_1_24_2","DOI":"10.1109\/CVPR52688.2022.01553"},{"doi-asserted-by":"publisher","key":"e_1_3_1_25_2","DOI":"10.1007\/978-3-319-46448-0_45"},{"doi-asserted-by":"publisher","key":"e_1_3_1_26_2","DOI":"10.1109\/TPAMI.2019.2957464"},{"doi-asserted-by":"publisher","key":"e_1_3_1_27_2","DOI":"10.1109\/ICCV51070.2023.00881"},{"key":"e_1_3_1_28_2","volume-title":"9th International Conference on Learning Representations (ICLR \u201921), Virtual Event","author":"Ke Guolin","year":"2021","unstructured":"Guolin Ke, Di He, and Tie-Yan Liu. 2021. Rethinking positional encoding in language pre-training. In 9th International Conference on Learning Representations (ICLR \u201921), Virtual Event. OpenReview.net. Retrieved from https:\/\/openreview.net\/forum?id=09-528y2Fgf"},{"key":"e_1_3_1_29_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations (ICLR \u201915)","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the 3rd International Conference on Learning Representations (ICLR \u201915)."},{"key":"e_1_3_1_30_2","first-page":"19","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923)","author":"Kou Yutong","year":"2023","unstructured":"Yutong Kou, Jin Gao, Bing Li, Gang Wang, Weiming Hu, Yizheng Wang, and Liang Li. 2023. ZoomTrack: Target-aware non-uniform resizing for efficient visual tracking. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NIPS \u201923). Curran Associates Inc., Red Hook, NY, Article 2217, 19 pages."},{"doi-asserted-by":"publisher","key":"e_1_3_1_31_2","DOI":"10.1109\/CVPR.2019.00441"},{"key":"e_1_3_1_32_2","first-page":", 12","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS \u201922)","author":"Lin Liting","year":"2022","unstructured":"Liting Lin, Heng Fan, Zhipeng Zhang, Yong Xu, and Haibin Ling. 2022. SwinTrack: A simple and strong baseline for transformer tracking. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NIPS \u201922). Curran Associates Inc., Red Hook, NY, Article 1218, 12 pages."},{"key":"e_1_3_1_33_2","volume-title":"Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022 (NeurIPS \u201922)","author":"Lin Liting","year":"2022","unstructured":"Liting Lin, Heng Fan, Zhipeng Zhang, Yong Xu, and Haibin Ling. 2022. SwinTrack: A simple and strong baseline for transformer tracking. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022 (NeurIPS \u201922). Retrieved from http:\/\/papers.nips.cc\/paper_files\/paper\/2022\/hash\/6a5c23219f401f3efd322579002dbb80-Abstract-Conference.html"},{"doi-asserted-by":"publisher","key":"e_1_3_1_34_2","DOI":"10.1109\/ICCV.2017.324"},{"doi-asserted-by":"publisher","key":"e_1_3_1_35_2","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_1_36_2","first-page":"13424","volume-title":"2021 IEEE\/CVF International Conference on Computer Vision","author":"Mayer Christoph","unstructured":"Christoph Mayer, Martin Danelljan, Danda Pani Paudel, and Luc Van Gool. [n.\u2009d.]. Learning target candidate association to keep track of what not to track. In 2021 IEEE\/CVF International Conference on Computer Vision. IEEE, 13424\u201313434."},{"doi-asserted-by":"publisher","key":"e_1_3_1_37_2","DOI":"10.1007\/978-3-030-01246-5_19"},{"key":"e_1_3_1_38_2","first-page":"817","volume-title":"33rd British Machine Vision Conference 2022 (BMVC \u201922)","author":"Noman Mubashir","year":"2022","unstructured":"Mubashir Noman, Wafa Al Ghallabi, Daniya Kareem, Christoph Mayer, Akshay Dudhane, Martin Danelljan, Hisham Cholakkal, Salman Khan, Luc Van Gool, and Fahad Shahbaz Khan. 2022. AVisT: A benchmark for visual object tracking in adverse visibility. In 33rd British Machine Vision Conference 2022 (BMVC \u201922). BMVA Press, 817. Retrieved from https:\/\/bmvc2022.mpi-inf.mpg.de\/817\/"},{"doi-asserted-by":"publisher","key":"e_1_3_1_39_2","DOI":"10.1109\/CVPR.2019.00075"},{"doi-asserted-by":"publisher","key":"e_1_3_1_40_2","DOI":"10.5555\/AAI28719412"},{"doi-asserted-by":"publisher","key":"e_1_3_1_41_2","DOI":"10.1109\/SLT54892.2023.10023097"},{"doi-asserted-by":"publisher","key":"e_1_3_1_42_2","DOI":"10.1609\/AAAI.V38I5.28286"},{"doi-asserted-by":"publisher","key":"e_1_3_1_43_2","DOI":"10.1609\/aaai.v37i2.25327"},{"doi-asserted-by":"publisher","key":"e_1_3_1_44_2","DOI":"10.1109\/CVPR52688.2022.00859"},{"key":"e_1_3_1_45_2","first-page":"5998","article-title":"Attention is all you need","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, 5998\u20136008.","journal-title":"Advances in Neural Information Processing Systems"},{"doi-asserted-by":"publisher","key":"e_1_3_1_46_2","DOI":"10.1109\/CVPR46437.2021.00162"},{"doi-asserted-by":"publisher","key":"e_1_3_1_47_2","DOI":"10.1109\/TIP.2024.3453028"},{"doi-asserted-by":"publisher","key":"e_1_3_1_48_2","DOI":"10.1109\/CVPR46437.2021.01355"},{"doi-asserted-by":"publisher","key":"e_1_3_1_49_2","DOI":"10.1109\/CVPR52729.2023.00935"},{"doi-asserted-by":"publisher","key":"e_1_3_1_50_2","DOI":"10.1109\/ICCV48922.2021.00988"},{"doi-asserted-by":"publisher","key":"e_1_3_1_51_2","DOI":"10.1007\/978-3-031-19803-8_43"},{"doi-asserted-by":"publisher","key":"e_1_3_1_52_2","DOI":"10.1109\/ICCV48922.2021.01028"},{"doi-asserted-by":"publisher","key":"e_1_3_1_53_2","DOI":"10.1007\/978-3-031-20047-2_20"},{"doi-asserted-by":"publisher","key":"e_1_3_1_54_2","DOI":"10.1007\/978-3-031-20047-2_20"},{"doi-asserted-by":"publisher","key":"e_1_3_1_55_2","DOI":"10.1109\/ICME55011.2023.00234"},{"doi-asserted-by":"publisher","key":"e_1_3_1_56_2","DOI":"10.1109\/WACV57701.2024.00634"},{"doi-asserted-by":"publisher","key":"e_1_3_1_57_2","DOI":"10.1145\/982452.982467"},{"doi-asserted-by":"publisher","key":"e_1_3_1_58_2","DOI":"10.1109\/ICME55011.2023.00247"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3732785","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,22]],"date-time":"2025-07-22T23:23:40Z","timestamp":1753226620000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3732785"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,22]]},"references-count":57,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,8,31]]}},"alternative-id":["10.1145\/3732785"],"URL":"https:\/\/doi.org\/10.1145\/3732785","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"type":"print","value":"2157-6904"},{"type":"electronic","value":"2157-6912"}],"subject":[],"published":{"date-parts":[[2025,7,22]]},"assertion":[{"value":"2024-04-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-20","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}