{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T15:59:39Z","timestamp":1781798379536,"version":"3.54.5"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2025,3,10]],"date-time":"2025-03-10T00:00:00Z","timestamp":1741564800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,4,30]]},"abstract":"<jats:p>Skeleton-based action recognition is beneficial for understanding human behavior in videos, and thus has received much attention in recent years as an important research area in action recognition. Current research focuses on designing more advanced algorithms to better extract spatio-temporal information from skeleton data. However, due to the small amount of data in the existing skeleton dataset and the lack of effective data augmentation methods, it is easy to lead to overfitting in model training. To address this challenge, we propose a mix-based data augmentation method, Joint Mixing Data Augmentation (JMDA), which can generally improve the effectiveness and robustness of various skeleton-based action recognition algorithms. In terms of spatial information, we introduce SpatialMix (SM), a method that projects the original 3D skeleton discrete information into a 2D space. Then, SM mixes the projected spatial information between two random samples during the training process to achieve the spatial-based mixing data augmentation. Concerning temporal information, we propose TemporalMix (TM). Leveraging the temporal continuity in skeleton data, we perform a temporal resize operation on the original skeleton data, and then merge two random samples during training to achieve the temporal-based mixed data augmentation. Additionally, we analyze the Feature Mismatch (FM) problem caused by introducing mix-based data augmentation into skeleton data. Then we propose a new data preprocessing method called Feature Alignment (FA) to effectively address this problem and improve model performance. Moreover, we propose a novel training pipeline, Joint Training Strategy (JTS), which combines multiple mix-based data augmentation methods for further improvement of model performance. Specifically, our proposed JMDA is plug-and-play and widely applicable to skeleton-based action recognition models. At the same time, the application of JMDA does not increase the model parameters and there is almost no additional training cost. We conduct extensive experiments on NTU RGB+D 60 and NTU RGB+D 120 datasets to demonstrate the effectiveness and robustness of the proposed JMDA on several mainstream skeleton-based action recognition algorithms.<\/jats:p>","DOI":"10.1145\/3700878","type":"journal-article","created":{"date-parts":[[2024,11,5]],"date-time":"2024-11-05T16:38:18Z","timestamp":1730824698000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Joint Mixing Data Augmentation for Skeleton-Based Action Recognition"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6738-223X","authenticated-orcid":false,"given":"Linhua","family":"Xiang","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, Hefei, China and National Engineering Research Centre of Speech and Language Information Processing, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1859-900X","authenticated-orcid":false,"given":"Zengfu","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China and National Engineering Research Centre of Speech and Language Information Processing, Hefei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,3,10]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2960588"},{"key":"e_1_3_1_3_2","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1109\/SIBGRAPI.2019.00011","volume-title":"Proceedings of the 32nd SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI \u201919)","author":"Caetano Carlos","year":"2019","unstructured":"Carlos Caetano, Fran\u00e7ois Br\u00e9mond, and William Robson Schwartz. 2019. Skeleton image representation for 3d action recognition based on tree structure and reference joints. In Proceedings of the 32nd SIBGRAPI Conference on Graphics, Patterns and Images (SIBGRAPI \u201919). IEEE, 16\u201323."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01311"},{"key":"e_1_3_1_5_2","first-page":"536","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"Cheng Ke","year":"2020","unstructured":"Ke Cheng, Yifan Zhang, Congqi Cao, Lei Shi, Jian Cheng, and Hanqing Lu. 2020. Decoupling GCN with dropgraph module for skeleton-based action recognition. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920). Springer, 536\u2013553."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00026"},{"key":"e_1_3_1_7_2","first-page":"20186","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chi Hyung-gun","year":"2022","unstructured":"Hyung-gun Chi, Myoung Hoon Ha, Seunggeun Chi, Sang Wan Lee, Qixing Huang, and Karthik Ramani. 2022. InfoGCN: Representation learning for human skeleton-based action recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 20186\u201320196."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01324"},{"key":"e_1_3_1_9_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_1_10_2","first-page":"1110","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Du Yong","year":"2015","unstructured":"Yong Du, Wei Wang, and Liang Wang. 2015. Hierarchical recurrent neural network for skeleton based action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1110\u20131118."},{"key":"e_1_3_1_11_2","unstructured":"Priya Goyal Piotr Doll\u00e1r Ross Girshick Pieter Noordhuis Lukasz Wesolowski Aapo Kyrola Andrew Tulloch Yangqing Jia and Kaiming He. 2017. Accurate large minibatch SGD: Training imagenet in 1 hour. arXiv:1706.02677. Retrieved from https:\/\/arxiv.org\/abs\/1706.02677"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3358415"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2941267"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12235"},{"key":"e_1_3_1_15_2","first-page":"8230","volume-title":"International Conference on Machine Learning","author":"Han Xiaotian","year":"2022","unstructured":"Xiaotian Han, Zhimeng Jiang, Ninghao Liu, and Xia Hu. 2022. G-mixup: Graph data augmentation for graph classification. In International Conference on Machine Learning. PMLR, 8230\u20138248."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01628"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2016.2628339"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1109\/ICeND.2014.6991357","volume-title":"Proceedings of the 3rd International Conference on e-Technologies and Networks for Development (ICeND \u201914)","author":"Htike Kyaw Kyaw","year":"2014","unstructured":"Kyaw Kyaw Htike, Othman O. Khalifa, Huda Adibah Mohd Ramli, and Mohammad A. M. Abushariah. 2014. Human activity recognition for video surveillance using sequences of postures. In Proceedings of the 3rd International Conference on e-Technologies and Networks for Development (ICeND \u201914). IEEE, 79\u201382."},{"key":"e_1_3_1_19_2","unstructured":"Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv:1312.6114."},{"key":"e_1_3_1_20_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"25","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 25."},{"key":"e_1_3_1_21_2","first-page":"10255","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Lee Jungho","year":"2023","unstructured":"Jungho Lee, Minhyeok Lee, Suhwan Cho, Sungmin Woo, Sungjun Jang, and Sangyoun Lee. 2023. Leveraging spatio-temporal dependency for skeleton-based action recognition. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 10255\u201310264."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3061115"},{"key":"e_1_3_1_23_2","first-page":"2737","volume-title":"Proceedings of the IEEE International Symposium on Circuits and Systems (ISCAS \u201908)","author":"Lin Weiyao","year":"2008","unstructured":"Weiyao Lin, Ming-Ting Sun, Radha Poovandran, and Zhengyou Zhang. 2008. Human activity recognition for video surveillance. In Proceedings of the IEEE International Symposium on Circuits and Systems (ISCAS \u201908). IEEE, 2737\u20132740."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3240472"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2916873"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.391"},{"key":"e_1_3_1_27_2","first-page":"441","volume-title":"Proceedings of the 17th European Conference on Computer Vision (ECCV \u201922)","author":"Liu Zicheng","year":"2022","unstructured":"Zicheng Liu, Siyuan Li, Di Wu, Zihan Liu, Zhiyuan Chen, Lirong Wu, and Stan Z. Li. 2022. Automix: Unveiling the power of mixup for stronger classifiers. In Proceedings of the 17th European Conference on Computer Vision (ECCV \u201922). Springer, 441\u2013458."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-60639-8_40"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3390127"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2021.103219"},{"key":"e_1_3_1_31_2","unstructured":"Jie Qin Jiemin Fang Qian Zhang Wenyu Liu Xingang Wang and Xinggang Wang. 2020. Resizemix: Mixing data with preserved object information and true labels. arXiv:2012.11101. Retrieved from https:\/\/arxiv.org\/abs\/2012.11101"},{"key":"e_1_3_1_32_2","doi-asserted-by":"crossref","unstructured":"Helei Qiu Biao Hou Bo Ren and Xiaohua Zhang. 2022. Spatio-temporal tuples transformer for skeleton-based action recognition. arXiv:2201.02849.","DOI":"10.1016\/j.neucom.2022.10.084"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2008.2005605"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.115"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01230"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2024.3359045"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.11212"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2670313"},{"issue":"4","key":"e_1_3_1_39_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.4018\/IJACI.2017100101","article-title":"Approaches and applications of virtual reality and gesture recognition: A review","volume":"8","author":"Sudha M. R.","year":"2017","unstructured":"M. R. Sudha, K. Sriraghav, Shomona Gracia Jacob, and S. Manisha. 2017. Approaches and applications of virtual reality and gesture recognition: A review. International Journal of Ambient Computing and Intelligence (IJACI) 8, 4 (2017), 1\u201318.","journal-title":"International Journal of Ambient Computing and Intelligence (IJACI)"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3117124"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3472722"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298664"},{"issue":"11","key":"e_1_3_1_43_2","article-title":"Visualizing data using t-SNE","volume":"9","author":"Maaten Laurens Van der","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, 11 (2008).","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_44_2","volume-title":"Proceedings of the Advances in Neural Information Processing Systems","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 30."},{"key":"e_1_3_1_45_2","first-page":"6438","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Verma Vikas","year":"2019","unstructured":"Vikas Verma, Alex Lamb, Christopher Beckham, Amir Najafi, Ioannis Mitliagkas, David Lopez-Paz, and Yoshua Bengio. 2019. Manifold mixup: Better representations by interpolating hidden states. In Proceedings of the International Conference on Machine Learning. PMLR, 6438\u20136447."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.387"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00544"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2869751"},{"key":"e_1_3_1_49_2","unstructured":"Qingtian Wang Jianlin Peng Shuze Shi Tingxi Liu Jiabin He and Renliang Weng. 2021. IIP-transformer: Intra-inter-part transformer for skeleton-based action recognition. arXiv:2110.13385. Retrieved from https:\/\/arxiv.org\/abs\/2110.13385"},{"key":"e_1_3_1_50_2","first-page":"6162","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"38","author":"Wu Xinyi","year":"2024","unstructured":"Xinyi Wu, Wentao Ma, Dan Guo, Tongqing Zhou, Shan Zhao, and Zhiping Cai. 2024. Text-based occluded person re-identification via multi-granularity contrastive consistency learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 6162\u20136170."},{"key":"e_1_3_1_51_2","unstructured":"Ziang Xie Sida I. Wang Jiwei Li Daniel L\u00e9vy Aiming Nie Dan Jurafsky and Andrew Y. Ng. 2017. Data noising as smoothing in neural network language models. arXiv:1703.02573."},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611900"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i3.20191"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3338533.3366569"},{"key":"e_1_3_1_56_2","unstructured":"Suorong Yang Weikang Xiao Mengcheng Zhang Suhan Guo Jian Zhao and Furao Shen. 2022. Image data augmentation for deep learning: A survey. arXiv:2204.08610. Retrieved from https:\/\/arxiv.org\/abs\/2204.08610"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00612"},{"key":"e_1_3_1_58_2","unstructured":"Hongyi Zhang Moustapha Cisse Yann N. Dauphin and David Lopez-Paz. 2017. mixup: Beyond empirical risk minimization. arXiv:1710.09412. Retrieved from https:\/\/arxiv.org\/abs\/1710.09412"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.233"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00119"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2016.03.014"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475473"},{"key":"e_1_3_1_63_2","first-page":"10608","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhou Huanyu","year":"2023","unstructured":"Huanyu Zhou, Qingjie Liu, and Yunhong Wang. 2023. Learning discriminative representations for skeleton based action recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 10608\u201310617."},{"key":"e_1_3_1_64_2","unstructured":"Kaiyang Zhou Yongxin Yang Yu Qiao and Tao Xiang. 2021. Domain generalization with mixstyle. arXiv:2104.02008. Retrieved from https:\/\/arxiv.org\/abs\/2104.02008"},{"key":"e_1_3_1_65_2","unstructured":"Yuxuan Zhou Chao Li Zhi-Qi Cheng Yifeng Geng Xuansong Xie and Margret Keuper. 2022. Hypergraph transformer for skeleton-based action recognition. arXiv:2211.09590. Retrieved from https:\/\/arxiv.org\/abs\/2211.09590"},{"key":"e_1_3_1_66_2","first-page":"15085","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhu Wentao","year":"2023","unstructured":"Wentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu, Wayne Wu, and Yizhou Wang. 2023. Motionbert: A unified perspective on learning human motion representations. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 15085\u201315099."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3700878","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3700878","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:10:23Z","timestamp":1750295423000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3700878"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,10]]},"references-count":65,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,4,30]]}},"alternative-id":["10.1145\/3700878"],"URL":"https:\/\/doi.org\/10.1145\/3700878","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,10]]},"assertion":[{"value":"2024-03-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}