{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T04:05:47Z","timestamp":1760241947531,"version":"build-2065373602"},"reference-count":46,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2018,10,29]],"date-time":"2018-10-29T00:00:00Z","timestamp":1540771200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"the National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61471154"],"award-info":[{"award-number":["61471154"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Anhui Province science and technology project","award":["1704d0802181"],"award-info":[{"award-number":["1704d0802181"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Video-based person re-identification is an important task with the challenges of lighting variation, low-resolution images, background clutter, occlusion, and human appearance similarity in the multi-camera visual sensor networks. In this paper, we propose a video-based person re-identification method called the end-to-end learning architecture with hybrid deep appearance-temporal feature. It can learn the appearance features of pivotal frames, the temporal features, and the independent distance metric of different features. This architecture consists of two-stream deep feature structure and two Siamese networks. For the first-stream structure, we propose the Two-branch Appearance Feature (TAF) sub-structure to obtain the appearance information of persons, and used one of the two Siamese networks to learn the similarity of appearance features of a pairwise person. To utilize the temporal information, we designed the second-stream structure that consisting of the Optical flow Temporal Feature (OTF) sub-structure and another Siamese network, to learn the person\u2019s temporal features and the distances of pairwise features. In addition, we select the pivotal frames of video as inputs to the Inception-V3 network on the Two-branch Appearance Feature sub-structure, and employ the salience-learning fusion layer to fuse the learned global and local appearance features. Extensive experimental results on the PRID2011, iLIDS-VID, and Motion Analysis and Re-identification Set (MARS) datasets showed that the respective proposed architectures reached 79%, 59% and 72% at Rank-1 and had advantages over state-of-the-art algorithms. Meanwhile, it also improved the feature representation ability of persons.<\/jats:p>","DOI":"10.3390\/s18113669","type":"journal-article","created":{"date-parts":[[2018,10,29]],"date-time":"2018-10-29T11:10:41Z","timestamp":1540811441000},"page":"3669","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Video-Based Person Re-Identification by an End-To-End Learning Architecture with Hybrid Deep Appearance-Temporal Feature"],"prefix":"10.3390","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1547-161X","authenticated-orcid":false,"given":"Rui","family":"Sun","sequence":"first","affiliation":[{"name":"School of Computer Science and Information Engineering, Hefei University of Technology, Feicui Road 420, Hefei 230000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiheng","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Engineering, Hefei University of Technology, Feicui Road 420, Hefei 230000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Miaomiao","family":"Xia","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Engineering, Hefei University of Technology, Feicui Road 420, Hefei 230000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Engineering, Hefei University of Technology, Feicui Road 420, Hefei 230000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2018,10,29]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"23509","DOI":"10.3390\/s141223509","article-title":"Generic learning-based ensemble framework for small sample size face recognition in multi-camera networks","volume":"14","author":"Zheng","year":"2014","journal-title":"Sensors"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"2501","DOI":"10.1109\/TPAMI.2016.2522418","article-title":"Person re-identification by discriminative selection in video ranking","volume":"38","author":"Wang","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Liu, K., Ma, B., Zhang, W., and Huang, R. (2015, January 7\u201313). A spatio-temporal appearance representation for video-based pedestrian re-identification. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.434"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"You, J., Wu, A., Li, X., and Zheng, W.-S. (2016, January 27\u201330). Top-push video-based person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.150"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"McLaughlin, N., del Rincon, J.M., and Miller, P. (2016, January 27\u201330). Recurrent convolutional network for video-based person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.148"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1629","DOI":"10.1109\/TPAMI.2014.2369055","article-title":"Person re-identificaiton by iterative re-weighted sparse ranking","volume":"37","author":"Lisanti","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Schumann, A., and Stiefelhagen, R. (2017, January 21\u201326). Person re-identification by deep learning attribute-complementary information. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.186"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Matsukawa, T., and Suzuki, E. (2016, January 4\u20138). Person re-identification using cnn features learned from combination of attributes. Proceedings of the International Conference on Pattern Recognition, Cancun, Mexico.","DOI":"10.1109\/ICPR.2016.7900000"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2768","DOI":"10.1109\/TCSVT.2017.2718188","article-title":"Learning bidirectional temporal cues for video-based person re-identification","volume":"28","author":"Zhang","year":"2017","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Chen, L., Yang, H., Zhu, J., Zhou, Q., Wu, S., and Gao, Z. (2017, January 21\u201326). Deep spatial-temporal fusion network for video-based person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.191"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Chung, D., Tahboub, K., and Delp, E.J. (2017, January 22\u201329). A two stream siamese convolutional neural network for person re-identification. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.218"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Hirzer, M., Beleznai, C., Roth, P.M., and Bischof, H. (2011, January 23\u201325). Person reidentification by descriptive and discriminative classification. Proceedings of the Scandinavian Conference on Image Analysis, Ystad, Sweden.","DOI":"10.1007\/978-3-642-21227-7_9"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, T., Gong, S., Zhu, X., and Wang, S. (2014, January 6\u201312). Person re-identification by video ranking. Proceedings of the European Conference on Computer Vision, Switzerland, Zurich.","DOI":"10.1007\/978-3-319-10593-2_45"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zheng, L., Bie, Z., Sun, Y., Wang, J., Su, C., Wang, S., and Tian, Q. (2016, January 11\u201314). Mars: A video benchmark for large-scale person re-identification. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46466-4_52"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Gray, D., and Tao, H. (2008, January 12\u201318). Viewpoint invariant pedestrian recognition with an ensemble of localized features. Proceedings of the 10th European Conference on Computer Vision, Marseille, France.","DOI":"10.1007\/978-3-540-88682-2_21"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhao, R., Ouyang, W., and Wang, X. (2013, January 23\u201328). Unsupervised salience learning for person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.460"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yang, Y., Yang, J., Yan, J., Liao, S., Yi, D., and Li, S.Z. (2014, January 6\u201312). Salient color names for person re-identification. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10590-1_35"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Matsukawa, T., Okabe, T., Suzuki, E., and Sato, Y. (2016, January 27\u201330). Hierarchical gaussian descriptor for person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.152"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Liao, S., Hu, Y., Zhu, X., and Li, S.Z. (2015, January 7\u201312). Person re-identification by local maximal occurrence representation and metric learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298832"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Kostinger, M., Hirzer, M., Wohlhart, P., Roth, P.M., and Bischof, H. (2012, January 16\u201321). Large scale metric learning from equivalence constraints. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6247939"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Li, Y., Wu, Z., Karanam, S., and Radke, R. (2015, January 7\u201310). Multi-shot human re-identification using adaptive fisher discriminant analysis. Proceedings of the British Machine Vision Conference, Swansea, UK.","DOI":"10.5244\/C.29.73"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Xu, S., Cheng, Y., Gu, K., Yang, Y., Chang, S., and Zhou, P. (2017, January 22\u201329). Jointly attentive spatial-temporal pooling networks for video-based person reidentification. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.507"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"717","DOI":"10.1109\/TIFS.2017.2765524","article-title":"Image to video person re-identification by learning heterogeneous dictionary pair with feature projection matrix","volume":"13","author":"Zhu","year":"2017","journal-title":"IEEE Trans. Inf. Forensics Secur."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Yi, D., Lei, Z., Liao, S., and Li, S.Z. (2014, January 24\u201328). Deep metric learning for person re-identification. Proceedings of the International Conference on Pattern Recognition, Stockholm, Sweden.","DOI":"10.1109\/ICPR.2014.16"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Cheng, D., Gong, Y., Zhou, S., Wang, J., and Zheng, N. (2016, January 27\u201330). Person re-identification by multi-channel parts-based cnn with improved triplet loss function. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.149"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"3492","DOI":"10.1109\/TIP.2017.2700762","article-title":"End-to-end comparative attention network for person re-identification","volume":"26","author":"Liu","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Chen, W., Chen, X., Zhang, J., and Huang, K. (2017, January 21\u201326). Beyond triplet loss: A deep quadruplet network for person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.145"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Li, W., Zhao, R., Xiao, T., and Wang, X. (2014, January 23\u201328). Deepreid: Deep filter pairing neural network for person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.27"},{"key":"ref_29","first-page":"1","article-title":"Compact appearance learning for video-based person re-identification","volume":"99","author":"Zhang","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_30","first-page":"191","article-title":"Electrical activity of muscles of the trunk during walking","volume":"111","author":"Waters","year":"1972","journal-title":"J. Anat."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 27\u201330). Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_32","unstructured":"Simonyan, K., and Zisserman, A. (2015, January 7\u20139). Very deep convolutional networks for large-scale image recognition. Proceedings of the International Conference on Learning Representations, San Diego, CA, USA."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_34","unstructured":"Lucas, B.D., and Kanade, T. (1981, January 24\u201328). An iterative image registration technique with an application to stereo vision. Proceedings of the International Joint Conference on Artificial Intelligence, Vancouver, BC, Canada."},{"key":"ref_35","first-page":"25","article-title":"Signature verification using a \u201csiamese\u201d time delay neural network","volume":"7","author":"Bromley","year":"1994","journal-title":"Int. J. Pattern Recognit. Artif. Intell."},{"key":"ref_36","unstructured":"Sun, Y., Chen, Y., Wang, X., and Tang, X. (2014, January 8\u201313). Deep learning face representation by joint identification-verification. Proceedings of the 27th International Conference on Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zheng, L., Shen, L., Tian, L., Wang, S., Wang, J., and Tian, Q. (2015, January 7\u201313). Scalable person re-identification: A benchmark. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.133"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object detection with discriminatively trained part-based models","volume":"32","author":"Felzenszwalb","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Dehghan, A., Assari, S.M., and Shah, M. (2015, January 7\u201312). GMMCP tracker: Globally optimal generalized maximum multi clique problem for multiple object tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299036"},{"key":"ref_40","unstructured":"Bolle, R.M., Connell, J.H., Pankanti, S., Ratha, N.K., and Senior, A.W. (2005, January 17\u201318). The relation between the roc curve and the cmc. Proceedings of the Fourth IEEE Workshop on Automatic Identification Advanced Technologies, Buffalo, NY, USA."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., and Darrell, T. (2014, January 3\u20137). Caffe: Convolutional architecture for fast feature embedding. Proceedings of the 22nd ACM international conference on Multimedia, Orlando, FL, USA.","DOI":"10.1145\/2647868.2654889"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Karanam, S., Li, Y., and Radke, R.J. (2015, January 7\u201313). Person re-identification with discriminatively trained viewpoint invariant dictionaries. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.513"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"998","DOI":"10.1109\/LSP.2016.2574323","article-title":"Person re-identification by exploiting spatio-temporal cues and multi-view metric learning","volume":"23","author":"Chen","year":"2016","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Yan, Y., Ni, B., Song, Z., Ma, C., Yan, Y., and Yang, X. (2016, January 8\u201316). Person re-identification via recurrent feature aggregation. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46466-4_42"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Zhou, Z., Huang, Y., Wang, W., Wang, L., and Tan, T. (2017, January 21\u201326). See the forest for the trees: Joint spatial and temporal recurrent neural networks for video-based person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.717"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Li, S., Bak, S., Carr, P., and Wang, X. (2018, January 18\u201322). Diversity regularized spatiotemporal attention for video-based person re-identification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00046"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/11\/3669\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T15:26:45Z","timestamp":1760196405000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/11\/3669"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,10,29]]},"references-count":46,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2018,11]]}},"alternative-id":["s18113669"],"URL":"https:\/\/doi.org\/10.3390\/s18113669","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2018,10,29]]}}}