{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T01:31:11Z","timestamp":1760232671258,"version":"build-2065373602"},"reference-count":46,"publisher":"MDPI AG","issue":"22","license":[{"start":{"date-parts":[[2022,11,20]],"date-time":"2022-11-20T00:00:00Z","timestamp":1668902400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Contrastive learning has received increasing attention in the field of skeleton-based action representations in recent years. Most contrastive learning methods use simple augmentation strategies to construct pairs of positive samples. When using such pairs of positive samples to learn action representations, deeper feature information cannot be learned, thus affecting the performance of downstream tasks. To solve the problem of insufficient learning ability, we propose an asymmetric data augmentation strategy and attempt to apply it to the training of 3D skeleton-based action representations. First, we carefully study the different characteristics presented by different skeleton views and choose a specific augmentation method for a certain view. Second, specific augmentation methods are incorporated into the left and right branches of the asymmetric data augmentation pipeline to increase the convergence difficulty of the contrastive learning task, thereby significantly improving the quality of the learned action representations. Finally, since many methods directly act on the joint view, the augmented samples are quite different from the original samples. We use random probability activation to transform the joint view to avoid extreme augmentation of the joint view. Extensive experiments on NTU RGB + D datasets show that our method is effective.<\/jats:p>","DOI":"10.3390\/s22228989","type":"journal-article","created":{"date-parts":[[2022,11,21]],"date-time":"2022-11-21T04:39:59Z","timestamp":1669005599000},"page":"8989","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Self-Supervised Action Representation Learning Based on Asymmetric Skeleton Data Augmentation"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4088-7473","authenticated-orcid":false,"given":"Hualing","family":"Zhou","sequence":"first","affiliation":[{"name":"College of Information Science and Engineering, Hunan Normal University, Changsha 410081, China"},{"name":"Key Laboratory of Sports Intelligence Reasearch, Hunan Normal University, Changsha 410081, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xi","family":"Li","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Hunan Normal University, Changsha 410081, China"},{"name":"Key Laboratory of Sports Intelligence Reasearch, Hunan Normal University, Changsha 410081, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dahong","family":"Xu","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Hunan Normal University, Changsha 410081, China"},{"name":"Key Laboratory of Sports Intelligence Reasearch, Hunan Normal University, Changsha 410081, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hong","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Hunan Normal University, Changsha 410081, China"},{"name":"Key Laboratory of Sports Intelligence Reasearch, Hunan Normal University, Changsha 410081, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianping","family":"Guo","sequence":"additional","affiliation":[{"name":"Key Laboratory of Sports Intelligence Reasearch, Hunan Normal University, Changsha 410081, China"},{"name":"College of Physical Culture, Hunan Normal University, Changsha 410081, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yihan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Sports Intelligence Reasearch, Hunan Normal University, Changsha 410081, China"},{"name":"College of Physical Culture, Hunan Normal University, Changsha 410081, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,11,20]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1109\/MMUL.2012.24","article-title":"Microsoft kinect sensor and its effect","volume":"19","author":"Zhang","year":"2012","journal-title":"IEEE Multimed."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Cao, Z., Hidalgo, G., Simon, T., Wei, S.E., and Sheikh, Y. (2018). OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields. arXiv.","DOI":"10.1109\/CVPR.2017.143"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Fang, H.S., Xie, S., Tai, Y.W., and Lu, C. (2017, January 22\u201329). RMPE: Regional multi-person pose estimation. Proceedings of the IEEE International Conference on Computer Vision(ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.256"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Xu, J., Yu, Z., Ni, B., Yang, J., Yang, X., and Zhang, W. (2020, January 14\u201319). Deep kinematics analysis for monocular 3d human pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR42600.2020.00098"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Ke, Q., Bennamoun, M., An, S., Sohel, F., and Boussaid, F. (2017, January 21\u201326). A new representation of skeleton sequences for 3d action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.486"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., and Lin, D. (2018, January 2\u20137). Spatial temporal graph convolutional networks for skeleton-based action recognition. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Si, C., Chen, W., Wang, W., Wang, L., and Tan, T. (2019, January 16\u201320). An attention enhanced graph convolutional lstm network for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00132"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Liu, Z., Zhang, H., Chen, Z., Wang, Z., and Ouyang, W. (2020, January 14\u201319). Disentangling and unifying graph convolutions for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR42600.2020.00022"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Chen, Z., Li, S., Yang, B., Li, Q., and Liu, H. (2021, January 2\u20139). Multi-scale spatial temporal graph convolutional network for skeleton-based action recognition. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual.","DOI":"10.1609\/aaai.v35i2.16197"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zheng, N., Wen, J., Liu, R., Long, L., Dai, J., and Gong, Z. (2018, January 2\u20137). Unsupervised representation learning with long-term dynamics for skeleton based action recognition. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11853"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Su, K., Liu, X., and Shlizerman, E. (2020, January 14\u201319). Predict & cluster: Unsupervised skeleton based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR42600.2020.00965"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Kundu, J.N., Gor, M., Uppala, P.K., and Radhakrishnan, V.B. (2019, January 7\u201311). Unsupervised feature learning of human actions as trajectories in pose embedding manifold. Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA.","DOI":"10.1109\/WACV.2019.00160"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Tian, Y., Krishnan, D., and Isola, P. (2020, January 23\u201328). Contrastive multiview coding. Proceedings of the European Conference on Computer Vision(ECCV), Virtual.","DOI":"10.1007\/978-3-030-58621-8_45"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Lin, L., Song, S., Yang, W., and Liu, J. (2020, January 12\u201316). Ms2l: Multi-task self-supervised learning for skeleton based action recognition. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3413548"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"90","DOI":"10.1016\/j.ins.2021.04.023","article-title":"Augmented skeleton based contrastive action learning with momentum lstm for unsupervised action recognition","volume":"569","author":"Rao","year":"2021","journal-title":"Inf. Sci."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Li, L., Wang, M., Ni, B., Wang, H., Yang, J., and Zhang, W. (2021, January 19\u201325). 3D human action representation learning via Cross-View consistency pursuit. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.00471"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020, January 14\u201319). Momentum contrast for unsupervised visual representation learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"ref_18","unstructured":"Chen, X., Fan, H., Girshick, R., and He, K. (2020). Improved baselines with momentum contrastive learning. arXiv."},{"key":"ref_19","unstructured":"Tian, Y., Sun, C., Poole, B., Krishnan, D., Schmid, C., and Isola, P. (2020). What makes for good views for contrastive learning?. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wang, X., and Qi, G.J. (2022). Contrastive Learning with Stronger Augmentations, IEEE.","DOI":"10.1109\/TPAMI.2022.3203630"},{"key":"ref_21","unstructured":"Wang, J., Liu, Z., Wu, Y., and Yuan, J. (2012, January 18\u201320). Mining actionlet ensemble for action recognition with depth cameras. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Providence, RI, USA."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Vemulapalli, R., Arrate, F., and Chellappa, R. (2014, January 23\u201328). Human action recognition by representing 3d skeletons as points in a lie group. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.82"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Vemulapalli, R., and Chellapa, R. (2016, January 27\u201330). Rolling rotationsfor recognizing human actions from 3d skeletal data. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.484"},{"key":"ref_24","unstructured":"Du, Y., Wang, W., and Wang, L. (2015, January 7\u201312). Hierarchical recurrent neural network for skeleton based action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"3459","DOI":"10.1109\/TIP.2018.2818328","article-title":"Spatio-temporal attention-based lstm networks for 3d action recognition and detection","volume":"27","author":"Song","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1963","DOI":"10.1109\/TPAMI.2019.2896631","article-title":"View adaptive neural networks for high performance skeleton-based human action recognition","volume":"41","author":"Zhang","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_27","unstructured":"Kremer, S.C., and Kolen, J.F. (2001). Gradientflow in Recurrent Nets: The Difficulty of Learning Long-Term Dependencies, IEEE Press. A Field Guide to Dynamical Recurrent Networks."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Du, Y., Fu, Y., and Wang, L. (2015, January 3\u20136). Skeleton based action recognition with convolutional neural network. Proceedings of the Asian Conference on Pattern Recognition (ACPR), Kuala Lumpur, Malaysia.","DOI":"10.1109\/ACPR.2015.7486569"},{"key":"ref_29","unstructured":"Li, C., Zhong, Q., Xie, D., and Pu, S. (2017, January 10\u201314). Skeleton-based action recognition with convolutional neural networks. Proceedings of the 2017 IEEE International Conference on Multimedia & Expo Workshops (ICMEW), Hong Kong, China."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2019, January 16\u201320). Two-stream adaptive graph convolutional networks for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01230"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Zhang, X., Xu, C., and Tao, D. (2020, January 14\u201319). Context aware graph convolution for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR42600.2020.01434"},{"key":"ref_32","unstructured":"Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. (2020, January 12\u201318). A simple framework for contrastive learning of visual representations. Proceedings of the International Conference on Machine Learning (ICML), Virtual."},{"key":"ref_33","unstructured":"Grill, J.B., Strub, F., Altch\u00e9, F., Tallec, C., Richemond, P.H., and Buchatskaya, E. (2020, January 6\u201312). Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning. Proceedings of the Neural Information Processing Systems (NeurIPS), Virtual."},{"key":"ref_34","unstructured":"Li, Y., Hu, P., Liu, Z., Peng, D., Zhou, J.T., and Peng, X. (2021, January 2\u20139). Contrastive Clustering. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021, January 11\u201317). Emerging properties in self-supervised vision transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Virtual.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Chen, X., and He, K. (2021, January 19\u201325). Exploring simple siamese representation learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.01549"},{"key":"ref_37","unstructured":"Guo, T., Liu, H., Chen, Z., Liu, M., Wang, T., and Ding, R. (March, January 22). Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-supervised Action Recognition. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Thoker, F.M., Doughty, H., and Snoek, C.G. (2021, January 18\u201320). Skeleton-contrastive 3D action representation learning. Proceedings of the 29th ACM International Conference on Multimedia, Virtual.","DOI":"10.1145\/3474085.3475307"},{"key":"ref_39","unstructured":"Oord, A.V.D., Li, Y., and Vinyals, O. (2018). Representation learning with contrastive predictive coding. arXiv."},{"key":"ref_40","unstructured":"Hinton, G., Vinyals, O., and Dean, J. (2015). Distilling the knowledge in a neural network. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2019, January 16\u201320). Skeleton-based action recognition with directed graph neural networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00810"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s40537-019-0197-0","article-title":"A survey on image data augmentation for deep learning","volume":"6","author":"Shorten","year":"2019","journal-title":"J. Big Data"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Shahroudy, A., Liu, J., Ng, T.T., and Wang, G. (2016, January 27\u201330). Ntu rgb+ d: A large scale dataset for 3d human activity analysis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.115"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"2684","DOI":"10.1109\/TPAMI.2019.2916873","article-title":"Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding","volume":"42","author":"Liu","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_45","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., and Chintala, S. (2019, January 8\u201314). Pytorch: An imperative style, high-performance deep learning library. Proceedings of the Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Yang, S., Liu, J., Lu, S., Er, M.H., and Kot, A.C. (2021, January 11\u201317). Skeleton cloud colorization for unsupervised 3D action representation learning. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Virtual.","DOI":"10.1109\/ICCV48922.2021.01317"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/22\/8989\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:22:25Z","timestamp":1760145745000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/22\/8989"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,11,20]]},"references-count":46,"journal-issue":{"issue":"22","published-online":{"date-parts":[[2022,11]]}},"alternative-id":["s22228989"],"URL":"https:\/\/doi.org\/10.3390\/s22228989","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2022,11,20]]}}}