{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T05:05:54Z","timestamp":1787029554849,"version":"3.56.0"},"reference-count":46,"publisher":"MDPI AG","issue":"19","license":[{"start":{"date-parts":[[2022,9,21]],"date-time":"2022-09-21T00:00:00Z","timestamp":1663718400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"European Regional Development Fund","award":["KK.01.1.1.01.0009"],"award-info":[{"award-number":["KK.01.1.1.01.0009"]}]},{"name":"European Regional Development Fund","award":["uniri-tehnic-18-295"],"award-info":[{"award-number":["uniri-tehnic-18-295"]}]},{"name":"University of Rijeka","award":["KK.01.1.1.01.0009"],"award-info":[{"award-number":["KK.01.1.1.01.0009"]}]},{"name":"University of Rijeka","award":["uniri-tehnic-18-295"],"award-info":[{"award-number":["uniri-tehnic-18-295"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Gait is a unique biometric trait with several useful properties. It can be recognized remotely and without the cooperation of the individual, with low-resolution cameras, and it is difficult to obscure. Therefore, it is suitable for crime investigation, surveillance, and access control. Existing approaches for gait recognition generally belong to the supervised learning domain, where all samples in the dataset are annotated. In the real world, annotation is often expensive and time-consuming. Moreover, convolutional neural networks (CNNs) have dominated the field of gait recognition for many years and have been extensively researched, while other recent methods such as vision transformer (ViT) remain unexplored. In this manuscript, we propose a self-supervised learning (SSL) approach for pretraining the feature extractor using the DINO model to automatically learn useful gait features with the vision transformer architecture. The feature extractor is then used for extracting gait features on which the fully connected neural network classifier is trained using the supervised approach. Experiments on CASIA-B and OU-MVLP gait datasets show the effectiveness of the proposed approach.<\/jats:p>","DOI":"10.3390\/s22197140","type":"journal-article","created":{"date-parts":[[2022,9,22]],"date-time":"2022-09-22T23:07:55Z","timestamp":1663888075000},"page":"7140","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":21,"title":["Gait Recognition with Self-Supervised Learning of Gait Features Based on Vision Transformers"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2431-9035","authenticated-orcid":false,"given":"Domagoj","family":"Pin\u010di\u0107","sequence":"first","affiliation":[{"name":"Faculty of Engineering, University of Rijeka, Vukovarska 58, 51000 Rijeka, Croatia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1026-7345","authenticated-orcid":false,"given":"Diego","family":"Su\u0161anj","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, University of Rijeka, Vukovarska 58, 51000 Rijeka, Croatia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0201-4177","authenticated-orcid":false,"given":"Kristijan","family":"Lenac","sequence":"additional","affiliation":[{"name":"Faculty of Engineering, University of Rijeka, Vukovarska 58, 51000 Rijeka, Croatia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,9,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"316","DOI":"10.1109\/TPAMI.2006.38","article-title":"Individual recognition using gait energy image","volume":"28","author":"Han","year":"2005","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"2164","DOI":"10.1109\/TPAMI.2011.260","article-title":"Human identification using temporal information preserving gait template","volume":"34","author":"Wang","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","unstructured":"Chao, H., He, Y., Zhang, J., and Feng, J. (February, January 27). Gaitset: Regarding gait as a set for cross-view gait recognition. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Fan, C., Peng, Y., Cao, C., Liu, X., Hou, S., Chi, J., Huang, Y., Li, Q., and He, Z. (2020, January 14\u201319). Gaitpart: Temporal part-based model for gait recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01423"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"107069","DOI":"10.1016\/j.patcog.2019.107069","article-title":"A model-based gait recognition method with body pose and human prior knowledge","volume":"98","author":"Liao","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_6","first-page":"84","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"209","DOI":"10.1109\/TPAMI.2016.2545669","article-title":"A comprehensive study on cross-view gait based human identification with deep cnns","volume":"39","author":"Wu","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1001","DOI":"10.1109\/TIP.2019.2926208","article-title":"Cross-view gait recognition by discriminative feature learning","volume":"29","author":"Zhang","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"22653","DOI":"10.1007\/s11042-020-09003-4","article-title":"Deep mutual learning network for gait recognition","volume":"79","author":"Wang","year":"2020","journal-title":"Multimed. Tools Appl."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"9767","DOI":"10.1109\/TGRS.2019.2929096","article-title":"Radar-based human gait recognition using dual-channel deep convolutional neural network","volume":"57","author":"Bai","year":"2019","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"669","DOI":"10.1109\/LGRS.2018.2806940","article-title":"Personnel recognition and gait classification based on multistatic micro-Doppler signatures using deep convolutional neural networks","volume":"15","author":"Chen","year":"2018","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Lenac, K., Su\u0161anj, D., Ramaki\u0107, A., and Pin\u010di\u0107, D. (2019). Extending appearance based gait recognition with depth data. Appl. Sci., 9.","DOI":"10.3390\/app9245529"},{"key":"ref_13","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Chen, Z., Xie, L., Niu, J., Liu, X., Wei, L., and Tian, Q. (2021, January 10\u201317). Visformer: The vision-friendly transformer. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00063"},{"key":"ref_16","first-page":"124","article-title":"View-invariant gait recognition with attentive recurrent learning of partial representations","volume":"3","author":"Etemad","year":"2020","journal-title":"IEEE Trans. Biom. Behav. Identity Sci."},{"key":"ref_17","first-page":"21271","article-title":"Bootstrap your own latent-a new approach to self-supervised learning","volume":"33","author":"Grill","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_18","unstructured":"Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. (2020, January 13\u201318). A simple framework for contrastive learning of visual representations. Proceedings of the International Conference on Machine Learning, PMLR, Virtual Event."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020, January 13\u201319). Momentum contrast for unsupervised visual representation learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021, January 10\u201317). Emerging properties in self-supervised vision transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"101539","DOI":"10.1016\/j.media.2019.101539","article-title":"Self-supervised learning for medical image analysis using image context restoration","volume":"58","author":"Chen","year":"2019","journal-title":"Med. Image Anal."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Tang, Y., Yang, D., Li, W., Roth, H.R., Landman, B., Xu, D., Nath, V., and Hatamizadeh, A. (2022, January 19\u201320). Self-supervised pre-training of swin transformers for 3d medical image analysis. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.02007"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"2052","DOI":"10.1016\/j.patrec.2010.05.027","article-title":"Gait recognition without subject cooperation","volume":"31","author":"Bashir","year":"2010","journal-title":"Pattern Recognit. Lett."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"2050266","DOI":"10.1142\/S0218126620502667","article-title":"Depth-based real-time gait recognition","volume":"29","author":"Lenac","year":"2020","journal-title":"J. Circuits Syst. Comput."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Shiraga, K., Makihara, Y., Muramatsu, D., Echigo, T., and Yagi, Y. (2016, January 13\u201316). Geinet: View-invariant gait recognition using a convolutional neural network. Proceedings of the 2016 International Conference on Biometrics (ICB), Halmstad, Sweden.","DOI":"10.1109\/ICB.2016.7550060"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"228","DOI":"10.1016\/j.patcog.2019.04.023","article-title":"A comprehensive study on gait biometrics using a joint CNN-based method","volume":"93","author":"Zhang","year":"2019","journal-title":"Pattern Recognit."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1016\/j.neucom.2021.04.081","article-title":"A novel view synthesis approach based on view space covering for gait recognition","volume":"453","author":"Liao","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"106988","DOI":"10.1016\/j.patcog.2019.106988","article-title":"Gaitnet: An end-to-end network for gait based human identification","volume":"96","author":"Song","year":"2019","journal-title":"Pattern Recognit."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1915","DOI":"10.1007\/s00371-021-02254-8","article-title":"Beyond view transformation: Feature distribution consistent GANs for cross-view gait recognition","volume":"38","author":"Wang","year":"2022","journal-title":"Vis. Comput."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"2023","DOI":"10.1007\/s10489-021-02484-2","article-title":"mmGaitSet: Multimodal based gait recognition for countering carrying and clothing changes","volume":"52","author":"Zhao","year":"2022","journal-title":"Appl. Intell."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"497","DOI":"10.1007\/s10044-020-00935-z","article-title":"Simple and efficient pose-based gait recognition method for challenging environments","volume":"24","author":"Lima","year":"2021","journal-title":"Pattern Anal. Appl."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Liao, R., Cao, C., Garcia, E.B., Yu, S., and Huang, Y. (2017, January 28\u201329). Pose-based temporal-spatial network (PTSN) for gait recognition with carrying and clothing variations. Proceedings of the Chinese Conference on Biometric Recognition, Shenzhen, China.","DOI":"10.1007\/978-3-319-69923-3_51"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wolf, T., Babaee, M., and Rigoll, G. (2016, January 25\u201328). Multi-view gait recognition using 3D convolutional neural networks. Proceedings of the 2016 IEEE International Conference on Image Processing (ICIP), Phoenix, AZ, USA.","DOI":"10.1109\/ICIP.2016.7533144"},{"key":"ref_34","first-page":"9912","article-title":"Unsupervised learning of visual features by contrasting cluster assignments","volume":"33","author":"Caron","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Cosma, A., and Radoi, I.E. (2021). Wildgait: Learning gait representations from raw surveillance streams. Sensors, 21.","DOI":"10.3390\/s21248387"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Liu, Y., Zeng, Y., Pu, J., Shan, H., He, P., and Zhang, J. (2021, January 6\u201312). Selfgait: A spatiotemporal representation learning method for self-supervised gait recognition. Proceedings of the ICASSP 2021\u20142021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9413894"},{"key":"ref_37","unstructured":"Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and J\u00e9gou, H. (2021, January 3\u201314). Training data-efficient image transformers & distillation through attention. Proceedings of the International Conference on Machine Learning, PMLR, Virtual Event."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Wen, Y., Zhang, K., Li, Z., and Qiao, Y. (2016, January 11\u201314). A discriminative feature learning approach for deep face recognition. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46478-7_31"},{"key":"ref_39","first-page":"441","article-title":"A framework for evaluating the effect of view angle, clothing and carrying condition on gait recognition","volume":"Volume 4","author":"Yu","year":"2006","journal-title":"Proceedings of the 18th International Conference on Pattern Recognition (ICPR\u201906)"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1186\/s41074-018-0039-6","article-title":"Multi-view large population gait dataset and its performance evaluation for cross-view gait recognition","volume":"10","author":"Takemura","year":"2018","journal-title":"IPSJ Trans. Comput. Vis. Appl."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., Joulin, A., and Self-Supervised Vision Transformers with DINO (2022, August 11). GitHub Repository. Available online: https:\/\/github.com\/facebookresearch\/dino.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_43","unstructured":"Loshchilov, I., and Hutter, F. (2017). Decoupled weight decay regularization. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"75381","DOI":"10.1109\/ACCESS.2020.2986554","article-title":"Flexible gait recognition based on flow regulation of local features between key frames","volume":"8","author":"Huang","year":"2020","journal-title":"IEEE Access"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Zhang, S., Wang, Y., and Li, A. (2021, January 19\u201325). Cross-view gait recognition with deep universal linear embeddings. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00898"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Wu, Z., Xiong, Y., Yu, S.X., and Lin, D. (2018, January 18\u201323). Unsupervised feature learning via non-parametric instance discrimination. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00393"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/19\/7140\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:36:11Z","timestamp":1760142971000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/19\/7140"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,21]]},"references-count":46,"journal-issue":{"issue":"19","published-online":{"date-parts":[[2022,10]]}},"alternative-id":["s22197140"],"URL":"https:\/\/doi.org\/10.3390\/s22197140","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,21]]}}}