{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T11:08:47Z","timestamp":1767956927136,"version":"3.49.0"},"reference-count":96,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T00:00:00Z","timestamp":1767916800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Primary Research & Development Plan of Heilongjiang Province","award":["GA23A903"],"award-info":[{"award-number":["GA23A903"]}]},{"DOI":"10.13039\/501100012130","name":"Key Laboratory of Avionics System Integrated Technology","doi-asserted-by":"publisher","award":["202400550P6001"],"award-info":[{"award-number":["202400550P6001"]}],"id":[{"id":"10.13039\/501100012130","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Fundamental Research Funds for the Central Universities in China","award":["3072024XX0602"],"award-info":[{"award-number":["3072024XX0602"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Existing detection-based trackers exploit temporal contexts by updating appearance models or modeling target motion. However, the sequential one-shot integration of temporal priors risks amplifying error accumulation, as frame-level template matching restricts comprehensive spatiotemporal analysis. To address this, we propose SCT-Diff, a video-level framework that holistically estimates target trajectories. Specifically, SCT-Diff processes video clips globally via a diffusion model to incorporate bidirectional spatiotemporal awareness, where reverse diffusion steps progressively refine noisy trajectory proposals into optimal predictions. Crucially, SCT-Diff enables iterative correction of historical trajectory hypotheses by observing future contexts within a sliding time window. This closed-loop feedback from future frames preserves temporal consistency and breaks the error propagation chain under complex appearance variations. For joint modeling of appearance and motion dynamics, we formulate trajectories as unified discrete token sequences. The designed Mamba-based expert decoder bridges visual features with language-formulated trajectories, enabling lightweight yet coherent sequence modeling. Extensive experiments demonstrate SCT-Diff\u2019s superior efficiency and performance, achieving 75.4% AO on GOT-10k while maintaining real-time computational efficiency.<\/jats:p>","DOI":"10.3390\/jimaging12010038","type":"journal-article","created":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T09:25:02Z","timestamp":1767950702000},"page":"38","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["SCT-Diff: Seamless Contextual Tracking via Diffusion Trajectory"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5341-4122","authenticated-orcid":false,"given":"Guohao","family":"Nie","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology, Harbin Engineering University, 145 Nantong Street, Harbin 150000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xingmei","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Harbin Engineering University, 145 Nantong Street, Harbin 150000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Debin","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Harbin Engineering University, 145 Nantong Street, Harbin 150000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"He","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Harbin Engineering University, 145 Nantong Street, Harbin 150000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,1,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"80297","DOI":"10.1109\/ACCESS.2023.3298440","article-title":"Transformers in single object tracking: An experimental survey","volume":"11","author":"Kugarajeevan","year":"2023","journal-title":"IEEE Access"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1435","DOI":"10.1007\/s13042-024-02345-7","article-title":"Beyond traditional visual object tracking: A survey","volume":"16","author":"Abdelaziz","year":"2025","journal-title":"Int. J. Mach. Learn. Cybern."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Bertinetto, L., Valmadre, J., Henriques, J.F., Vedaldi, A., and Torr, P.H. (2016, January 11\u201314). Fully-convolutional siamese networks for object tracking. Proceedings of the Computer Vision\u2013ECCV 2016 Workshops, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-48881-3_56"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Chen, X., Yan, B., Zhu, J., Wang, D., Yang, X., and Lu, H. (2021, January 20\u201325). Transformer tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00803"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Li, B., Yan, J., Wu, W., Zhu, Z., and Hu, X. (2018, January 18\u201323). High performance visual tracking with siamese region proposal network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00935"},{"key":"ref_6","first-page":"16743","article-title":"Swintrack: A simple and strong baseline for transformer tracking","volume":"35","author":"Lin","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Yan, B., Peng, H., Fu, J., Wang, D., and Lu, H. (2021, January 10\u201317). Learning spatio-temporal transformer for visual tracking. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01028"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"144","DOI":"10.3390\/jimaging11050144","article-title":"Motion-Perception Multi-Object Tracking (MPMOT): Enhancing Multi-Object Tracking Performance via Motion-Aware Data Association and Trajectory Connection","volume":"11","author":"Meng","year":"2025","journal-title":"J. Imaging"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Cui, Y., Jiang, C., Wang, L., and Wu, G. (2022, January 18\u201324). Mixformer: End-to-end tracking with iterative mixed attention. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01324"},{"key":"ref_10","first-page":"2321","article-title":"Compact transformer tracker with correlative masked modeling","volume":"37","author":"Song","year":"2023","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Cai, W., Liu, Q., and Wang, Y. (2024, January 16\u201322). Hiptrack: Visual tracking with historical prompts. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01822"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Mayer, C., Danelljan, M., Bhat, G., Paul, M., Paudel, D.P., Yu, F., and Van Gool, L. (2022, January 18\u201324). Transforming model prediction for tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00853"},{"key":"ref_13","unstructured":"Zhang, L., Gonzalez-Garcia, A., Weijer, J.V.D., Danelljan, M., and Khan, F.S. (November, January 27). Learning the model update for siamese trackers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Xie, J., Zhong, B., Mo, Z., Zhang, S., Shi, L., Song, S., and Ji, R. (2024, January 16\u201322). Autoregressive queries for adaptive tracking with spatio-temporal transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01826"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Wei, X., Bai, Y., Zheng, Y., Shi, D., and Gong, Y. (2023, January 17\u201324). Autoregressive Visual Tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00935"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Bai, Y., Zhao, Z., Gong, Y., and Wei, X. (2024, January 16\u201322). Artrackv2: Prompting autoregressive tracker where to look and how to describe. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01802"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Chen, X., Peng, H., Wang, D., Lu, H., and Hu, H. (2023, January 17\u201324). Seqtrack: Sequence to sequence learning for visual object tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01400"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"4895","DOI":"10.1109\/TCSVT.2021.3056684","article-title":"Dynamic attention guided multi-trajectory analysis for single object tracking","volume":"31","author":"Wang","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"3425","DOI":"10.1109\/TCSVT.2022.3233636","article-title":"Trajectory guided robust visual object tracking with selective remedy","volume":"33","author":"Wang","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"8588","DOI":"10.1007\/s10489-021-02829-x","article-title":"Non-linear target trajectory prediction for robust visual tracking","volume":"52","author":"Xu","year":"2022","journal-title":"Appl. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"171","DOI":"10.3390\/jimaging10070171","article-title":"Deep Efficient Data Association for Multi-Object Tracking: Augmented with SSIM-Based Ambiguity Elimination","volume":"10","author":"Prasannakumar","year":"2024","journal-title":"J. Imaging"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Xie, F., Chu, L., Li, J., Lu, Y., and Ma, C. (2023, January 17\u201324). Videotrack: Learning to track objects via video transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02186"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ye, B., Chang, H., Ma, B., Shan, S., and Chen, X. (2022, January 23\u201327). Joint feature learning and relation modeling for tracking: A one-stream framework. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-20047-2_20"},{"key":"ref_24","first-page":"8780","article-title":"Diffusion models beat gans on image synthesis","volume":"34","author":"Dhariwal","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_25","unstructured":"Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. arXiv."},{"key":"ref_26","first-page":"4713","article-title":"Image super-resolution via iterative refinement","volume":"45","author":"Saharia","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Li, B., Wu, W., Wang, Q., Zhang, F., Xing, J., and Yan, J. (2019, January 15\u201320). Siamrpn++: Evolution of siamese visual tracking with very deep networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00441"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Gao, S., Zhou, C., and Zhang, J. (2023, January 17\u201324). Generalized relation modeling for transformer tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01792"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., and Girshick, R. (2022, January 18\u201324). Masked autoencoders are scalable vision learners. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"ref_30","first-page":"21002","article-title":"Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection","volume":"33","author":"Li","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_31","unstructured":"Zhou, X., Wang, D., and Kr\u00e4henb\u00fchl, P. (2019). Objects as points. arXiv."},{"key":"ref_32","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_33","unstructured":"Bhat, G., Danelljan, M., Gool, L.V., and Timofte, R. (November, January 27). Learning discriminative model prediction for tracking. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Danelljan, M., Bhat, G., Shahbaz Khan, F., and Felsberg, M. (2017, January 21\u201326). Eco: Efficient convolution operators for tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.733"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Danelljan, M., Bhat, G., Khan, F.S., and Felsberg, M. (2019, January 15\u201320). Atom: Accurate tracking by overlap maximization. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00479"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Lan, J.P., Cheng, Z.Q., He, J.Y., Li, C., Luo, B., Bao, X., Xiang, W., Geng, Y., and Xie, X. (2023, January 4\u201310). Procontext: Exploring progressive context transformer for tracking. Proceedings of the ICASSP 2023\u20132023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece.","DOI":"10.1109\/ICASSP49357.2023.10094971"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Li, X., Ma, C., Wu, B., He, Z., and Yang, M.H. (2019, January 15\u201320). Target-aware deep tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00146"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Wang, G., Luo, C., Sun, X., Xiong, Z., and Zeng, W. (2020, January 13\u201319). Tracking by instance detection: A meta-learning approach. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00632"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Yang, T., and Chan, A.B. (2018, January 8\u201314). Learning dynamic memory networks for object tracking. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01240-3_10"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Dai, K., Zhang, Y., Wang, D., Li, J., Lu, H., and Yang, X. (2020, January 13\u201319). High-performance long-term tracking with meta-updater. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00633"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"110504","DOI":"10.1016\/j.knosys.2023.110504","article-title":"Hierarchical memory-guided long-term tracking with meta transformer inquiry network","volume":"269","author":"Wang","year":"2023","journal-title":"Knowl.-Based Syst."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"112229","DOI":"10.1016\/j.asoc.2024.112229","article-title":"Temporal relation transformer for robust visual tracking with dual-memory learning","volume":"167","author":"Nie","year":"2024","journal-title":"Appl. Soft Comput."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"6897","DOI":"10.1109\/TCSVT.2023.3272319","article-title":"Memory network with pixel-level spatio-temporal learning for visual object tracking","volume":"33","author":"Zhou","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_44","first-page":"773","article-title":"Target-aware tracking with long-term context attention","volume":"37","author":"He","year":"2023","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_45","unstructured":"Oh, S.W., Lee, J.Y., Xu, N., and Kim, S.J. (November, January 27). Video object segmentation using space-time memory networks. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Wang, N., Zhou, W., Wang, J., and Li, H. (2021, January 20\u201325). Transformer meets tracker: Exploiting temporal context for robust visual tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00162"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Bhat, G., Danelljan, M., Van Gool, L., and Timofte, R. (2020, January 23\u201328). Know your surroundings: Exploiting scene information for object tracking. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58592-1_13"},{"key":"ref_48","unstructured":"Sauer, A., Aljalbout, E., and Haddadin, S. (2019). Tracking holistic object representations. arXiv."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Fu, Z., Liu, Q., Fu, Z., and Wang, Y. (2021, January 20\u201325). Stmtrack: Template-free visual tracking with space-time memory networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01356"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"1129","DOI":"10.1007\/s12596-023-01293-9","article-title":"Discriminative learning of online appearance modeling methods for visual tracking","volume":"53","author":"Liao","year":"2024","journal-title":"J. Opt."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"2125","DOI":"10.1109\/TCSVT.2023.3301933","article-title":"Toward unified token learning for vision-language tracking","volume":"34","author":"Zheng","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_52","first-page":"579","article-title":"Frido: Feature pyramid diffusion for complex scene image synthesis","volume":"37","author":"Fan","year":"2023","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K. (2023, January 17\u201324). Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"ref_54","first-page":"36479","article-title":"Photorealistic text-to-image diffusion models with deep language understanding","volume":"35","author":"Saharia","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_55","unstructured":"Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., and Gafni, O. (2022). Make-a-video: Text-to-video generation without text-video data. arXiv."},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Yang, R., Srivastava, P., and Mandt, S. (2023). Diffusion probabilistic modeling for video generation. Entropy, 25.","DOI":"10.3390\/e25101469"},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"4115","DOI":"10.1109\/TPAMI.2024.3355414","article-title":"Motiondiffuse: Text-driven human motion generation with diffusion model","volume":"46","author":"Zhang","year":"2024","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Huang, R., Zhao, Z., Liu, H., Liu, J., Cui, C., and Ren, Y. (2022, January 10\u201314). Prodiff: Progressive fast diffusion model for high-quality text-to-speech. Proceedings of the 30th ACM International Conference on Multimedia, Lisbon, Portugal.","DOI":"10.1145\/3503161.3547855"},{"key":"ref_59","unstructured":"Kim, S., Kim, H., and Yoon, S. (2022). Guided-tts 2: A diffusion model for high-quality adaptive text-to-speech with untranscribed data. arXiv."},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Levkovitch, A., Nachmani, E., and Wolf, L. (2022). Zero-shot voice conditioning for denoising diffusion tts models. arXiv.","DOI":"10.21437\/Interspeech.2022-10045"},{"key":"ref_61","unstructured":"Wu, S., and Shi, Z. (2026, January 05). Stochastic Differential Equation is All You Need for Voice Generation. Available online: https:\/\/papers.ssrn.com\/sol3\/papers.cfm?abstract_id=4156409."},{"key":"ref_62","doi-asserted-by":"crossref","first-page":"1720","DOI":"10.1109\/TASLP.2023.3268730","article-title":"Diffsound: Discrete diffusion model for text-to-sound generation","volume":"31","author":"Yang","year":"2023","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_63","first-page":"17981","article-title":"Structured denoising diffusion models in discrete state-spaces","volume":"34","author":"Austin","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_64","unstructured":"Gong, S., Li, M., Feng, J., Wu, Z., and Kong, L. (2022). Diffuseq: Sequence to sequence text generation with diffusion models. arXiv."},{"key":"ref_65","first-page":"4328","article-title":"Diffusion-lm improves controllable text generation","volume":"35","author":"Li","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Gong, S., Li, M., Feng, J., Wu, Z., and Kong, L. (2023). Diffuseq-v2: Bridging discrete and continuous text spaces for accelerated seq2seq diffusion models. arXiv.","DOI":"10.18653\/v1\/2023.findings-emnlp.660"},{"key":"ref_67","unstructured":"He, Y., Cai, Z., Gan, X., and Chang, B. (2023). DiffCap: Exploring continuous diffusion on image captioning. arXiv."},{"key":"ref_68","first-page":"3991","article-title":"Diffusiontrack: Diffusion model for multi-object tracking","volume":"38","author":"Luo","year":"2024","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Ji, Y., Chen, Z., Xie, E., Hong, L., Liu, X., Liu, Z., Lu, T., Li, Z., and Luo, P. (2023, January 1\u20136). Ddp: Diffusion model for dense visual prediction. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.01987"},{"key":"ref_70","doi-asserted-by":"crossref","unstructured":"Brempong, E.A., Kornblith, S., Chen, T., Parmar, N., Minderer, M., and Norouzi, M. (2022, January 18\u201324). Denoising pretraining for semantic segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPRW56347.2022.00462"},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Chen, T., Li, L., Saxena, S., Hinton, G., and Fleet, D.J. (2023, January 18\u201324). A generalist framework for panoptic segmentation of images and videos. Proceedings of the IEEE\/CVF International Conference on Computer Vision, New Orleans, LA, USA.","DOI":"10.1109\/ICCV51070.2023.00090"},{"key":"ref_72","first-page":"14715","article-title":"Diffusion models as plug-and-play priors","volume":"35","author":"Graikos","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_73","unstructured":"Kim, B., Oh, Y., and Ye, J.C. (2022). Diffusion adversarial representation learning for self-supervised vessel segmentation. arXiv."},{"key":"ref_74","unstructured":"Wolleb, J., Sandk\u00fchler, R., Bieder, F., Valmaggia, P., and Cattin, P.C. (2022, January 6\u20138). Diffusion models for implicit image segmentation ensembles. Proceedings of the International Conference on Medical Imaging with Deep Learning, Zurich, Switzerland."},{"key":"ref_75","doi-asserted-by":"crossref","unstructured":"Chen, S., Sun, P., Song, Y., and Luo, P. (2023, January 1\u20136). Diffusiondet: Diffusion model for object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.01816"},{"key":"ref_76","doi-asserted-by":"crossref","unstructured":"Xie, F., Wang, Z., and Ma, C. (2024, January 16\u201322). Diffusiontrack: Point set diffusion model for visual object tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01808"},{"key":"ref_77","unstructured":"Gu, A., and Dao, T. (2023). Mamba: Linear-time sequence modeling with selective state spaces. arXiv."},{"key":"ref_78","unstructured":"Chen, T., Zhang, R., and Hinton, G. (2022). Analog bits: Generating discrete data using diffusion models with self-conditioning. arXiv."},{"key":"ref_79","unstructured":"Nichol, A.Q., and Dhariwal, P. (2021, January 18\u201324). Improved denoising diffusion probabilistic models. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_80","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the Computer Vision\u2013ECCV 2014: 13th European Conference, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_81","doi-asserted-by":"crossref","unstructured":"Fan, H., Lin, L., Yang, F., Chu, P., Deng, G., Yu, S., Bai, H., Xu, Y., Liao, C., and Ling, H. (2019, January 15\u201320). Lasot: A high-quality benchmark for large-scale single object tracking. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00552"},{"key":"ref_82","doi-asserted-by":"crossref","first-page":"1562","DOI":"10.1109\/TPAMI.2019.2957464","article-title":"Got-10k: A large high-diversity benchmark for generic object tracking in the wild","volume":"43","author":"Huang","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_83","doi-asserted-by":"crossref","unstructured":"Muller, M., Bibi, A., Giancola, S., Alsubaihi, S., and Ghanem, B. (2018, January 8\u201314). Trackingnet: A large-scale dataset and benchmark for object tracking in the wild. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01246-5_19"},{"key":"ref_84","doi-asserted-by":"crossref","unstructured":"Wang, X., Shu, X., Zhang, Z., Jiang, B., Wang, Y., Tian, Y., and Wu, F. (2021, January 20\u201325). Towards more flexible and accurate object tracking with natural language: Algorithms and benchmark. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01355"},{"key":"ref_85","doi-asserted-by":"crossref","unstructured":"Wu, Y., Lim, J., and Yang, M.H. (2013, January 23\u201328). Online object tracking: A benchmark. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.312"},{"key":"ref_86","doi-asserted-by":"crossref","unstructured":"Kiani Galoogahi, H., Fagg, A., Huang, C., Ramanan, D., and Lucey, S. (2017, January 22\u201329). Need for speed: A benchmark for higher frame rate object tracking. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.128"},{"key":"ref_87","doi-asserted-by":"crossref","first-page":"5630","DOI":"10.1109\/TIP.2015.2482905","article-title":"Encoding color information for visual tracking: Algorithms and benchmark","volume":"24","author":"Liang","year":"2015","journal-title":"IEEE Trans. Image Process."},{"key":"ref_88","unstructured":"Song, J., Meng, C., and Ermon, S. (2020). Denoising diffusion implicit models. arXiv."},{"key":"ref_89","unstructured":"Loshchilov, I., and Hutter, F. (2017, January 24\u201326). Decoupled Weight Decay Regularization. Proceedings of the International Conference on Learning Representations, Toulon, France."},{"key":"ref_90","doi-asserted-by":"crossref","unstructured":"Nam, H., and Han, B. (2016, January 27\u201330). Learning Multi-Domain Convolutional Neural Networks for Visual Tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.465"},{"key":"ref_91","doi-asserted-by":"crossref","unstructured":"Gao, S., Zhou, C., Ma, C., Wang, X., and Yuan, J. (2022, January 23\u201327). Aiatrack: Attention in attention for transformer visual tracking. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-20047-2_9"},{"key":"ref_92","first-page":"4838","article-title":"Explicit Visual Prompts for Visual Object Tracking","volume":"38","author":"Shi","year":"2024","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_93","first-page":"7979","article-title":"MIMTrack: In-Context Tracking via Masked Image Modeling","volume":"39","author":"Wang","year":"2025","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_94","unstructured":"Li, K., Li, X., Wang, Y., He, Y., Wang, Y., Wang, L., and Qiao, Y. (October, January 29). Videomamba: State space model for efficient video understanding. Proceedings of the European Conference on Computer Vision, Milan, Italy."},{"key":"ref_95","doi-asserted-by":"crossref","first-page":"19413","DOI":"10.1109\/TITS.2025.3590935","article-title":"Signeye: Traffic sign interpretation from vehicle first-person view","volume":"26","author":"Yang","year":"2025","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_96","doi-asserted-by":"crossref","first-page":"805","DOI":"10.1007\/s11263-023-01912-9","article-title":"SignParser: An end-to-end framework for traffic sign understanding","volume":"132","author":"Guo","year":"2024","journal-title":"Int. J. Comput. Vis."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/1\/38\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T09:40:41Z","timestamp":1767951641000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/1\/38"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,9]]},"references-count":96,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1]]}},"alternative-id":["jimaging12010038"],"URL":"https:\/\/doi.org\/10.3390\/jimaging12010038","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,9]]}}}