{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T15:58:23Z","timestamp":1782575903294,"version":"3.54.5"},"reference-count":51,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2023,3,1]],"date-time":"2023-03-01T00:00:00Z","timestamp":1677628800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"CRC research","award":["PN-III 1\/2018"],"award-info":[{"award-number":["PN-III 1\/2018"]}]},{"name":"Google IoT\/Wearables Student Grants","award":["PN-III 1\/2018"],"award-info":[{"award-number":["PN-III 1\/2018"]}]},{"name":"Keysight Master Research Sponsorship","award":["PN-III 1\/2018"],"award-info":[{"award-number":["PN-III 1\/2018"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The manner of walking (gait) is a powerful biometric that is used as a unique fingerprinting method, allowing unobtrusive behavioral analytics to be performed at a distance without subject cooperation. As opposed to more traditional biometric authentication methods, gait analysis does not require explicit cooperation of the subject and can be performed in low-resolution settings, without requiring the subject\u2019s face to be unobstructed\/clearly visible. Most current approaches are developed in a controlled setting, with clean, gold-standard annotated data, which powered the development of neural architectures for recognition and classification. Only recently has gait analysis ventured into using more diverse, large-scale, and realistic datasets to pretrained networks in a self-supervised manner. Self-supervised training regime enables learning diverse and robust gait representations without expensive manual human annotations. Prompted by the ubiquitous use of the transformer model in all areas of deep learning, including computer vision, in this work, we explore the use of five different vision transformer architectures directly applied to self-supervised gait recognition. We adapt and pretrain the simple ViT, CaiT, CrossFormer, Token2Token, and TwinsSVT on two different large-scale gait datasets: GREW and DenseGait. We provide extensive results for zero-shot and fine-tuning on two benchmark gait recognition datasets, CASIA-B and FVG, and explore the relationship between the amount of spatial and temporal gait information used by the visual transformer. Our results show that in designing transformer models for processing motion, using a hierarchical approach (i.e., CrossFormer models) on finer-grained movement fairs comparatively better than previous whole-skeleton approaches.<\/jats:p>","DOI":"10.3390\/s23052680","type":"journal-article","created":{"date-parts":[[2023,3,1]],"date-time":"2023-03-01T02:02:56Z","timestamp":1677636176000},"page":"2680","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["Exploring Self-Supervised Vision Transformers for Gait Recognition in the Wild"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0307-2520","authenticated-orcid":false,"given":"Adrian","family":"Cosma","sequence":"first","affiliation":[{"name":"Faculty of Automatic Control and Computer Science, University Politehnica of Bucharest, 006042 Bucharest, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5324-0428","authenticated-orcid":false,"given":"Andy","family":"Catruna","sequence":"additional","affiliation":[{"name":"Faculty of Automatic Control and Computer Science, University Politehnica of Bucharest, 006042 Bucharest, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1177-5288","authenticated-orcid":false,"given":"Emilian","family":"Radoi","sequence":"additional","affiliation":[{"name":"Faculty of Automatic Control and Computer Science, University Politehnica of Bucharest, 006042 Bucharest, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,3,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"e19555","DOI":"10.1097\/MD.0000000000019555","article-title":"Gait pattern analysis and clinical subgroup identification: A retrospective observational study","volume":"99","author":"Kyeong","year":"2020","journal-title":"Medicine"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"580","DOI":"10.1097\/PSY.0b013e3181a2515c","article-title":"Embodiment of Sadness and Depression\u2014Gait Patterns Associated With Dysphoric Mood","volume":"71","author":"Michalak","year":"2009","journal-title":"Psychosom. Med."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"330","DOI":"10.1249\/01.mss.0000247001.94470.21","article-title":"Gait-related risk factors for exercise-related lower-leg pain during shod running","volume":"39","author":"Willems","year":"2007","journal-title":"Med. Sci. Sports Exerc."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"70497","DOI":"10.1109\/ACCESS.2018.2879896","article-title":"Vision-based gait recognition: A survey","volume":"6","author":"Singh","year":"2018","journal-title":"IEEE Access"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Makihara, Y., Nixon, M.S., and Yagi, Y. (2020). Gait recognition: Databases, representations, and applications. Comput. Vis. Ref. Guide, 1\u201313.","DOI":"10.1007\/978-3-030-03243-2_883-1"},{"key":"ref_6","unstructured":"Yu, S., Tan, D., and Tan, T. (2006, January 20\u201324). A framework for evaluating the effect of view angle, clothing and carrying condition on gait recognition. Proceedings of the 18th International Conference on Pattern Recognition (ICPR\u201906), Hong Kong, China."},{"key":"ref_7","unstructured":"Zhu, Z., Guo, X., Yang, T., Huang, J., Deng, J., Huang, G., Du, D., Lu, J., and Zhou, J. (2021, January 11\u201317). Gait Recognition in the Wild: A Benchmark. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Montreal, BC, Canada."},{"key":"ref_8","unstructured":"Chao, H., He, Y., Zhang, J., and Feng, J. (February, January 27). Gaitset: Regarding gait as a set for cross-view gait recognition. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Fan, C., Peng, Y., Cao, C., Liu, X., Hou, S., Chi, J., Huang, Y., Li, Q., and He, Z. (2020, January 13\u201319). Gaitpart: Temporal part-based model for gait recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01423"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Cosma, A., and Radoi, I.E. (2021). WildGait: Learning Gait Representations from Raw Surveillance Streams. Sensors, 21.","DOI":"10.3390\/s21248387"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Catruna, A., Cosma, A., and Radoi, I.E. (2021, January 15\u201318). From Face to Gait: Weakly-Supervised Learning of Gender Information from Walking Patterns. Proceedings of the 2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021), Jodhpur, India.","DOI":"10.1109\/FG52635.2021.9666987"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Cosma, A., and Radoi, E. (2022). Learning Gait Representations with Noisy Multi-Task Learning. Sensors, 22.","DOI":"10.3390\/s22186803"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Kirkcaldy, B.D. (1985). Individual Differences in Movement, MTP Press Lancaster.","DOI":"10.1007\/978-94-009-4912-6"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zheng, J., Liu, X., Liu, W., He, L., Yan, C., and Mei, T. (2022, January 19\u201320). Gait Recognition in the Wild with Dense 3D Representations and A Benchmark. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01959"},{"key":"ref_15","first-page":"857","article-title":"Self-supervised learning: Generative or contrastive","volume":"35","author":"Liu","year":"2021","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_16","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019, January 2\u20137). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, MN, USA."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021, January 11\u201317). Emerging properties in self-supervised vision transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., and Lin, D. (2018, January 2\u20137). Spatial temporal graph convolutional networks for skeleton-based action recognition. Proceedings of the Thirty-second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Xu, C., Makihara, Y., Liao, R., Niitsuma, H., Li, X., Yagi, Y., and Lu, J. (2021, January 3\u20138). Real-Time Gait-Based Age Estimation and Gender Classification From a Single Image. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA.","DOI":"10.1109\/WACV48630.2021.00350"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"172","DOI":"10.1109\/TPAMI.2019.2929257","article-title":"OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields","volume":"43","author":"Cao","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Li, J., Wang, C., Zhu, H., Mao, Y., Fang, H.S., and Lu, C. (2019, January 15\u201320). Crowdpose: Efficient crowded scenes pose estimation and a new benchmark. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01112"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Liu, Z., Zhang, H., Chen, Z., Wang, Z., and Ouyang, W. (2020, January 13\u201319). Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00022"},{"key":"ref_24","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Touvron, H., Cord, M., Sablayrolles, A., Synnaeve, G., and J\u00e9gou, H. (2021, January 11\u201317). Going deeper with image transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00010"},{"key":"ref_26","unstructured":"Wang, W., Yao, L., Chen, L., Lin, B., Cai, D., He, X., and Liu, W. (2021). CrossFormer: A versatile vision transformer hinging on cross-scale attention. arXiv."},{"key":"ref_27","first-page":"9355","article-title":"Twins: Revisiting the design of spatial attention in vision transformers","volume":"34","author":"Chu","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z.H., Tay, F.E., Feng, J., and Yan, S. (2021, January 11\u201317). Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00060"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Tran, L., Yin, X., Atoum, Y., Wan, J., Wang, N., and Liu, X. (2019, January 15\u201320). Gait Recognition via Disentangled Representation Learning. Proceedings of the IEEE Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00484"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Teepe, T., Khan, A., Gilg, J., Herzog, F., H\u00f6rmann, S., and Rigoll, G. (2021, January 19\u201322). Gaitgraph: Graph convolutional network for skeleton-based gait recognition. Proceedings of the 2021 IEEE International Conference on Image Processing (ICIP), Anchorage, AK, USA.","DOI":"10.1109\/ICIP42928.2021.9506717"},{"key":"ref_31","unstructured":"Fu, Y., Wei, Y., Zhou, Y., Shi, H., Huang, G., Wang, X., Yao, Z., and Huang, T. (February, January 27). Horizontal Pyramid Matching for Person Re-Identification. Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Song, Y.F., Zhang, Z., Shan, C., and Wang, L. (2020, January 12\u201316). Stronger, faster and more explainable: A graph convolutional baseline for skeleton-based action recognition. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3413802"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Li, N., and Zhao, X. (2022). A Strong and Robust Skeleton-based Gait Recognition Method with Gait Periodicity Priors. IEEE Trans. Multimed.","DOI":"10.1109\/TMM.2022.3154609"},{"key":"ref_34","first-page":"6000","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L. (2021, January 11\u201317). Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00061"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_37","first-page":"30008","article-title":"Focal attention for long-range interactions in vision transformers","volume":"34","author":"Yang","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_38","unstructured":"Chen, C.F., Panda, R., and Fan, Q. (2021). Regionvit: Regional-to-local attention for vision transformers. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Graham, B., El-Nouby, A., Touvron, H., Stock, P., Joulin, A., J\u00e9gou, H., and Douze, M. (2021, January 11\u201317). Levit: A vision transformer in convnet\u2019s clothing for faster inference. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01204"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lu\u010di\u0107, M., and Schmid, C. (2021, January 11\u201317). Vivit: A video vision transformer. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"ref_41","unstructured":"Bertasius, G., Wang, H., and Torresani, L. (2021, January 18\u201324). Is space-time attention all you need for video understanding?. Proceedings of the International Conference on Machine Learning (ICML), Virtual."},{"key":"ref_42","unstructured":"Xu, X., Meng, Q., Qin, Y., Guo, J., Zhao, C., Zhou, F., and Lei, Z. (2021, January 2\u20139). Searching for alignment in face recognition. Proceedings of the AAAI Conference on Artificial Intelligence, Virtually."},{"key":"ref_43","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 6\u201311). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the International Conference on Machine Learning (PMLR), Lille, France."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"2405","DOI":"10.1109\/TCSVT.2018.2864148","article-title":"Action recognition with spatio\u2013temporal visual attention on skeleton image sequences","volume":"29","author":"Yang","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_45","first-page":"18661","article-title":"Supervised Contrastive Learning","volume":"Volume 33","author":"Larochelle","year":"2020","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_46","unstructured":"Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. (2020, January 13\u201318). A simple framework for contrastive learning of visual representations. Proceedings of the International Conference on Machine Learning (PMLR), Virtual."},{"key":"ref_47","unstructured":"Smith, L.N. (2015). No More Pesky Learning Rate Guessing Games. arXiv."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Wang, J., Jiao, J., and Liu, Y.H. (2020, January 23\u201327). Self-supervised video representation learning by pace prediction. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-030-58520-4_30"},{"key":"ref_49","first-page":"1","article-title":"The OU-ISIR Gait Database Comprising the Large Population Dataset with Age and Performance Evaluation of Age Estimation","volume":"9","author":"Xu","year":"2017","journal-title":"IPSJ Trans. Comput. Vis. Appl."},{"key":"ref_50","unstructured":"Zhang, T., Wu, F., Katiyar, A., Weinberger, K.Q., and Artzi, Y. (2020, January 26\u201330). Revisiting Few-sample BERT Fine-tuning. Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia."},{"key":"ref_51","first-page":"2579","article-title":"Visualizing Data using t-SNE","volume":"9","author":"Hinton","year":"2008","journal-title":"J. Mach. Learn. Res."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/5\/2680\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:44:51Z","timestamp":1760121891000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/5\/2680"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,1]]},"references-count":51,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["s23052680"],"URL":"https:\/\/doi.org\/10.3390\/s23052680","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,1]]}}}