{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,14]],"date-time":"2025-11-14T07:37:23Z","timestamp":1763105843879,"version":"build-2065373602"},"reference-count":56,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2021,12,15]],"date-time":"2021-12-15T00:00:00Z","timestamp":1639526400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61402540","60903222","61672538","61272024"],"award-info":[{"award-number":["61402540","60903222","61672538","61272024"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>A convolutional neural network can easily fall into local minima for insufficient data, and the needed training is unstable. Many current methods are used to solve these problems by adding pedestrian attributes, pedestrian postures, and other auxiliary information, but they require additional collection, which is time-consuming and laborious. Every video sequence frame has a different degree of similarity. In this paper, multi-level fusion temporal\u2013spatial co-attention is adopted to improve person re-identification (reID). For a small dataset, the improved network can better prevent over-fitting and reduce the dataset limit. Specifically, the concept of knowledge evolution is introduced into video-based person re-identification to improve the backbone residual neural network (ResNet). The global branch, local branch, and attention branch are used in parallel for feature extraction. Three high-level features are embedded in the metric learning network to improve the network\u2019s generalization ability and the accuracy of video-based person re-identification. Simulation experiments are implemented on small datasets PRID2011 and iLIDS-VID, and the improved network can better prevent over-fitting. Experiments are also implemented on MARS and DukeMTMC-VideoReID, and the proposed method can be used to extract more feature information and improve the network\u2019s generalization ability. The results show that our method achieves better performance. The model achieves 90.15% Rank1 and 81.91% mAP on MARS.<\/jats:p>","DOI":"10.3390\/e23121686","type":"journal-article","created":{"date-parts":[[2021,12,15]],"date-time":"2021-12-15T21:47:36Z","timestamp":1639604856000},"page":"1686","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Multi-Level Fusion Temporal\u2013Spatial Co-Attention for Video-Based Person Re-Identification"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0924-3079","authenticated-orcid":false,"given":"Shengyu","family":"Pei","sequence":"first","affiliation":[{"name":"School of Automation, Central South University, Changsha 410075, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0172-4070","authenticated-orcid":false,"given":"Xiaoping","family":"Fan","sequence":"additional","affiliation":[{"name":"School of Automation, Central South University, Changsha 410075, China"},{"name":"School of Information Technology and Management, Hunan University of Finance and Economics, Changsha 410205, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,12,15]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Zhou, K., Yang, Y., Cavallaro, A., and Xiang, T. Learning generalisable omni-scale representations for person re-identification. IEEE Trans. Pattern Anal. Mach. Intell., 2021. in press.","DOI":"10.1109\/TPAMI.2021.3069237"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1649","DOI":"10.1109\/TPAMI.2019.2954313","article-title":"Person re-identification with deep kronecker-product matching and group-shuffling random walk","volume":"43","author":"Shen","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","unstructured":"Miao, J., Wu, Y., and Yang, Y. (2021). Identifying visible parts via pose estimation for occluded person re-identification. IEEE Trans. Neural Networks Learn. Syst., 1\u201311."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1016\/j.neucom.2020.12.018","article-title":"Triplet online instance matching loss for person re-identification","volume":"433","author":"Li","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1460","DOI":"10.1109\/TPAMI.2020.2976969","article-title":"Ordered or orderless: A revisit for video based person re-identification","volume":"43","author":"Zhang","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"8821","DOI":"10.1109\/TIP.2020.3001693","article-title":"Adaptive graph representation learning for video person re-identification","volume":"29","author":"Wu","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"7679","DOI":"10.1007\/s10489-021-02271-z","article-title":"Image generation and constrained two-stage feature fusion for person re-identification","volume":"51","author":"Zhang","year":"2021","journal-title":"Appl. Intell."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"107799","DOI":"10.1016\/j.patcog.2020.107799","article-title":"3d-GAT: 3d-guided adversarial transform network for person re-identification in unseen domains","volume":"112","author":"Zhang","year":"2021","journal-title":"Pattern Recognit."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"013001","DOI":"10.1117\/1.JEI.30.1.013001","article-title":"Adaptive spatial scale person reidentification","volume":"30","author":"Pei","year":"2021","journal-title":"J. Electron. Imaging"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"107688","DOI":"10.1016\/j.patcog.2020.107688","article-title":"Hypergraph video pedestrian re-identification based on posture structure relationship and action constraints","volume":"111","author":"Hu","year":"2021","journal-title":"Pattern Recognit."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"788","DOI":"10.1007\/s10489-020-01844-8","article-title":"Discriminative feature extraction for video person re-identification via multi-task network","volume":"51","author":"Song","year":"2021","journal-title":"Appl. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"2788","DOI":"10.1109\/TCSVT.2017.2715499","article-title":"Video-based person re-identification with accumulative motion context","volume":"28","author":"Liu","year":"2017","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"176","DOI":"10.1016\/j.ins.2020.04.007","article-title":"Pose-guided spatiotemporal alignment for video-based person re-identification","volume":"527","author":"Gao","year":"2020","journal-title":"Inf. Sci."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"503","DOI":"10.1109\/TCSVT.2020.2988034","article-title":"Hierarchical temporal modeling with mutual distance matching for video based person re-identification","volume":"31","author":"Li","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Gao, C., Yao, R., Zhou, Y., Zhao, J., Fang, L., and Hu, F. (2021). Efficient lightweight video person re-identification with online difference discrimination module. Multimed. Tools Appl., 1\u201313.","DOI":"10.1007\/s11042-021-10543-6"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3402666","article-title":"Correlation discrepancy insight network for video re-identification","volume":"16","author":"Ruan","year":"2020","journal-title":"ACM Trans. Multimed. Comput. Commun. Appl. (TOMM)"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"e1964","DOI":"10.1002\/cav.1964","article-title":"One-shot video-based person re-identification with variance subsampling algorithm","volume":"31","author":"Zhao","year":"2020","journal-title":"Comput. Animat. Virtual Worlds"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"12841","DOI":"10.1007\/s00521-020-04730-z","article-title":"Scale-fusion framework for improving video-based person re-identification performance","volume":"32","author":"Cheng","year":"2020","journal-title":"Neural Comput. Appl."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Eom, C., Lee, G., Lee, J., and Ham, B. (2021, January 1\u20134). Video-based Person Re-identification with Spatial and Temporal Memory Networks. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Nashville, TN, USA.","DOI":"10.1109\/ICCV48922.2021.01182"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Hou, R., Chang, H., Ma, B., Huang, R., and Shan, S. (2021, January 1\u20134). BiCnet-TKS: Learning Efficient Spatial-Temporal Representation for Video Person Re-Identification. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00205"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Liu, C.T., Chen, J.C., Chen, C.S., and Chien, S.Y. (2021, January 1\u20134). Video-based Person Re-identification without Bells and Whistles. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPRW53098.2021.00165"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Wang, Y., Zhang, P., Gao, S., Geng, X., Lu, H., and Wang, D. (2021, January 1\u20134). Pyramid Spatial-Temporal Aggregation for Video-Based Person Re-Identification. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Nashville, TN, USA.","DOI":"10.1109\/ICCV48922.2021.01181"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Aich, A., Zheng, M., Karanam, S., Chen, T., Roy-Chowdhury, A.K., and Wu, Z. (2021, January 1\u20134). Spatio-temporal representation factorization for video-based person re-identification. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Nashville, TN, USA.","DOI":"10.1109\/ICCV48922.2021.00022"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Liu, J., Zha, Z.J., Wu, W., Zheng, K., and Sun, Q. (2021, January 1\u20134). Spatial-Temporal Correlation and Topology Learning for Person Re-Identification in Videos. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00435"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Taha, A., Shrivastava, A., and Davis, L.S. (2021, January 1\u20134). Knowledge evolution in neural networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01265"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zheng, L., Bie, Z., Sun, Y., Wang, J., Su, C., Wang, S., and Tian, Q. (2016, January 8\u201316). MARS: A video benchmark for large-scale person re-identification. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46466-4_52"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Wu, Y., Lin, Y., Dong, X., Yan, Y., Ouyang, W., and Yang, Y. (2018, January 18\u201323). Exploit the unknown gradually: One-shot video-based person re-identification by stepwise learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00543"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Hirzer, M., Beleznai, C., Roth, P.M., and Bischof, H. (2011, January 23\u201325). Person re-identification by descriptive and discriminative classification. Proceedings of the Scandinavian Conference on Image Analysis, Ystad, Sweden.","DOI":"10.1007\/978-3-642-21227-7_9"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Wang, T., Gong, S., Zhu, X., and Wang, S. (2014, January 6\u201312). Person re-identification by video ranking. Proceedings of the European conference on computer vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10593-2_45"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"2501","DOI":"10.1109\/TPAMI.2016.2522418","article-title":"Person re-identification by discriminative selection in video ranking","volume":"38","author":"Wang","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"197","DOI":"10.1016\/j.patcog.2016.11.018","article-title":"Person re-identification by unsupervised video matching","volume":"65","author":"Ma","year":"2017","journal-title":"Pattern Recognit."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Li, M., Zhu, X., and Gong, S. (2018, January 8\u201314). Unsupervised person re-identification by deep learning tracklet association. Proceedings of the European conference on computer vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01225-0_45"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Huang, Y., Wang, W., Wang, L., and Tan, T. (2017, January 21\u201326). See the forest for the trees: Joint spatial and temporal recurrent neural networks for video-based person re-identification. Proceedings of the IEEE conference on computer vision and pattern recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.717"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Liu, Y., Yan, J., and Ouyang, W. (2017, January 21\u201326). Quality aware network for set to set recognition. Proceedings of the IEEE conference on computer vision and pattern recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.499"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Li, D., Chen, X., Zhang, Z., and Huang, K. (2017, January 21\u201326). Learning deep context-aware features over body and latent parts for person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.782"},{"key":"ref_36","unstructured":"Hermans, A., Beyer, L., and Leibe, B. (2017). In defense of the triplet loss for person re-identification. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Song, C., Huang, Y., Ouyang, W., and Wang, L. (2018, January 18\u201323). Mask-guided contrastive attention model for person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00129"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Li, S., Bak, S., Carr, P., and Wang, X. (2018, January 18\u201323). Diversity regularized spatiotemporal attention for video-based person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00046"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Si, J., Zhang, H., Li, C.G., Kuen, J., Kong, X., Kot, A.C., and Wang, G. (2018, January 18\u201323). Dual attention matching network for context-aware feature sequence based person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00562"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Suh, Y., Wang, J., Tang, S., Mei, T., and Lee, K.M. (2018, January 8\u201314). Part-aligned bilinear representations for person re-identification. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_25"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Chen, D., Li, H., Xiao, T., Yi, S., and Wang, X. (2018, January 18\u201323). Video person re-identification with competitive snippet-similarity aggregation and co-attentive snippet embedding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00128"},{"key":"ref_42","unstructured":"Liu, Y., Yuan, Z., Zhou, W., and Li, H. (February, January 27). Spatial and temporal mutual promotion for video-based person re-identification. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_43","unstructured":"Li, J., Zhang, S., and Huang, T. (February, January 27). Multi-scale 3d convolution network for video based person re-identification. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_44","unstructured":"Fu, Y., Wang, X., Wei, Y., and Huang, T. (February, January 27). STA: Spatial-temporal attention for large-scale video-based person re-identification. Proceedings of the AAAI conference on artificial intelligence, Honolulu, HI, USA."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Li, J., Wang, J., Tian, Q., Gao, W., and Zhang, S. (2019, January 15\u201320). Global-local temporal representations for video person re-identification. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00406"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Pathak, P., Eshratifar, A.E., and Gormish, M. (2020, January 7\u201312). Video Person Re-ID: Fantastic Techniques and Where to Find Them. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i10.7219"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Yang, J., Zheng, W., Yang, Q., Chen, Y., and Tian, Q. (2020, January 13\u201319). Spatial-temporal graph convolutional network for video-based person re-identification. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00335"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"6266","DOI":"10.1109\/TIP.2021.3093759","article-title":"A Two-Stream Dynamic Pyramid Representation Model for Video-Based Person Re-Identification","volume":"30","author":"Yang","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Li, Y., Zhuo, L., Li, J., Zhang, J., Liang, X., and Tian, Q. (2017, January 21\u201326). Video-based person re-identification by deep feature guided pooling. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.188"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"McLaughlin, N., Del Rincon, J.M., and Miller, P. (2016, January 27\u201330). Recurrent convolutional network for video-based person re-identification. Proceedings of the IEEE conference on computer vision and pattern recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.148"},{"key":"ref_51","unstructured":"Wu, L., Shen, C., and Hengel, A.V.D. (2016). Deep recurrent convolutional networks for video-based person re-identification: An end-to-end approach. arXiv."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Yan, Y., Ni, B., Song, Z., Ma, C., Yan, Y., and Yang, X. (2016, January 8\u201316). Person re-identification via recurrent feature aggregation. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46466-4_42"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Wu, Y., Qiu, J., Takamatsu, J., and Ogasawara, T. (2018, January 2\u20137). Temporal-enhanced convolutional network for person re-identification. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12264"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Hou, R., Chang, H., Ma, B., Shan, S., and Chen, X. (2020, January 23\u201328). Temporal complementary learning for video person re-identification. Proceedings of the European conference on computer vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58595-2_24"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Liu, L., Yang, X., Wang, N., and Gao, X. (2021, January 20\u201325). Viewing from Frequency Domain: A DCT-based Information Enhancement Network for Video Person Re-Identification. Proceedings of the 29th ACM International Conference on Multimedia, Nashville, TN, USA.","DOI":"10.1145\/3474085.3475566"},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1016\/j.neucom.2021.10.018","article-title":"What-Where-When Attention Network for video-based person re-identification","volume":"468","author":"Zhang","year":"2022","journal-title":"Neurocomputing"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/12\/1686\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:49:01Z","timestamp":1760168941000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/12\/1686"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,15]]},"references-count":56,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2021,12]]}},"alternative-id":["e23121686"],"URL":"https:\/\/doi.org\/10.3390\/e23121686","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2021,12,15]]}}}