{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,18]],"date-time":"2026-02-18T23:52:04Z","timestamp":1771458724490,"version":"3.50.1"},"reference-count":59,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2023,12,8]],"date-time":"2023-12-08T00:00:00Z","timestamp":1701993600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2018YFC0407905"],"award-info":[{"award-number":["2018YFC0407905"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Human action recognition (HAR) as the most representative human-centred computer vision task is critical in human resource management (HRM), especially in human resource recruitment, performance appraisal, and employee training. Currently, prevailing approaches to human action recognition primarily emphasize either temporal or spatial features while overlooking the intricate interplay between these two dimensions. This oversight leads to less precise and robust action classification within complex human resource recruitment environments. In this paper, we propose a novel human action recognition methodology for human resource recruitment environments, which aims at symmetrically harnessing temporal and spatial information to enhance the performance of human action recognition. Specifically, we compute Depth Motion Maps (DMM) and Depth Temporal Maps (DTM) from depth video sequences as space and time descriptors, respectively. Subsequently, a novel feature fusion technique named Center Boundary Collaborative Canonical Correlation Analysis (CBCCCA) is designed to enhance the fusion of space and time features by collaboratively learning the center and boundary information of feature class space. We then introduce a spatio-temporal information filtration module to remove redundant information introduced by spatio-temporal fusion and retain discriminative details. Finally, a Support Vector Machine (SVM) is employed for human action recognition. Extensive experiments demonstrate that the proposed method has the ability to significantly improve human action recognition performance.<\/jats:p>","DOI":"10.3390\/sym15122177","type":"journal-article","created":{"date-parts":[[2023,12,8]],"date-time":"2023-12-08T03:03:33Z","timestamp":1702004613000},"page":"2177","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Spatio-Temporal Information Fusion and Filtration for Human Action Recognition"],"prefix":"10.3390","volume":"15","author":[{"given":"Man","family":"Zhang","sequence":"first","affiliation":[{"name":"College of Information Science and Technology & College of Artificial Intelligence, Nanjing Forestry University, Nanjing 210037, China"},{"name":"College of Social Sciences, University of Birmingham, Birmingham B15 2TT, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xing","family":"Li","sequence":"additional","affiliation":[{"name":"College of Information Science and Technology & College of Artificial Intelligence, Nanjing Forestry University, Nanjing 210037, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qianhan","family":"Wu","sequence":"additional","affiliation":[{"name":"School of Computer and Information, Hohai University, Nanjing 211100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,12,8]]},"reference":[{"key":"ref_1","first-page":"7607864","article-title":"Enterprise Human Resources Recruitment Management Model in the Era of Mobile Internet","volume":"2022","author":"Yang","year":"2022","journal-title":"Mob. Inf. Syst."},{"key":"ref_2","unstructured":"Tanti, L., Puspasari, R., and Triandi, B. (2018, January 7\u20139). Employee Performance Assessment with Profile Matching Method. Proceedings of the 2018 6th International Conference on Cyber and IT Service Management (CITSM), Parapat, Indonesia."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1913","DOI":"10.1057\/s41291-023-00234-5","article-title":"Sustainable training practices: Predicting job satisfaction and employee behavior using machine learning techniques","volume":"22","author":"Gupta","year":"2023","journal-title":"Asian Bus. Manag."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"3141","DOI":"10.1109\/TCSVT.2021.3103677","article-title":"FEXNet: Foreground Extraction Network for Human Action Recognition","volume":"32","author":"Shen","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1016\/j.neucom.2022.01.091","article-title":"Human action recognition by multiple spatial clues network","volume":"483","author":"Zheng","year":"2022","journal-title":"Neurocomputing"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"226","DOI":"10.1016\/j.engappai.2017.10.001","article-title":"Deep convolutional framework for abnormal behavior detection in a smart surveillance system","volume":"67","author":"Ko","year":"2018","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Rodomagoulakis, I., Kardaris, N., Pitsikalis, V., Mavroudi, E., Katsamanis, A., Tsiami, A., and Maragos, P. (2016, January 20\u201325). Multimodal human action recognition in assistive human-robot interaction. Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472168"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"103371","DOI":"10.1016\/j.jvcir.2021.103371","article-title":"Complex Network-based features extraction in RGB-D human action recognition","volume":"82","author":"Yang","year":"2022","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1109\/34.910878","article-title":"The recognition of human movement using temporal templates","volume":"23","author":"Bobick","year":"2001","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Yang, X., Zhang, C., and Tian, Y. (2012). Recognizing Actions Using Depth Motion Maps-Based Histograms of Oriented Gradients, Association for Computing Machinery.","DOI":"10.1145\/2393347.2396382"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"21085","DOI":"10.1007\/s11042-019-7365-2","article-title":"3D human action analysis and recognition through GLAC descriptor on 2D motion and static posture images","volume":"78","author":"Bulbul","year":"2019","journal-title":"Multimed. Tools Appl."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"5275","DOI":"10.1109\/TIP.2018.2855438","article-title":"Information Fusion for Human Action Recognition via Biset\/Multiset Globality Locality Preserving Canonical Correlation Analysis","volume":"27","author":"Elmadany","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"103716","DOI":"10.1016\/j.jvcir.2022.103716","article-title":"Spatial and temporal information fusion for human action recognition via Center Boundary Balancing Multimodal Classifier","volume":"90","author":"Li","year":"2023","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Li, W., Zhang, Z., and Liu, Z. (2010, January 13\u201318). Action recognition based on a bag of 3d points. Proceedings of the 2010 IEEE Computer Society Conference On Computer Vision and Pattern Recognition-Workshops, San Francisco, CA, USA.","DOI":"10.1109\/CVPRW.2010.5543273"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Chen, C., Jafari, R., and Kehtarnavaz, N. (2015, January 27\u201330). UTD-MHAD: A multimodal dataset for human action recognition utilizing a depth camera and a wearable inertial sensor. Proceedings of the 2015 IEEE International Conference on Image Processing (ICIP), Quebec City, QC, Canada.","DOI":"10.1109\/ICIP.2015.7350781"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Lin, Y.C., Hu, M.C., Cheng, W.H., Hsieh, Y.H., and Chen, H.M. (2012, January 5\u20138). Human action recognition and retrieval using sole depth information. Proceedings of the Acm International Conference on Multimedia, Hong Kong, China.","DOI":"10.1145\/2393347.2396381"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Chen, C., Jafari, R., and Kehtarnavaz, N. (2015, January 5\u20139). Action Recognition from Depth Sequences Using Depth Motion Maps-based Local Binary Patterns. Proceedings of the 2015 IEEE Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV.2015.150"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1806","DOI":"10.1109\/TSMC.2018.2850149","article-title":"Deep Convolutional Neural Networks for Human Action Recognition Using Depth Maps and Postures","volume":"49","author":"Kamel","year":"2019","journal-title":"IEEE Trans. Syst. Man Cybern.-Syst."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"24119","DOI":"10.1007\/s11042-022-12091-z","article-title":"3DFCNN: Real-time action recognition using 3D deep neural networks with raw depth information","volume":"81","author":"Sarker","year":"2022","journal-title":"Multimed. Tools Appl."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1850033","DOI":"10.1142\/S0218001418500337","article-title":"Multi-View Hierarchical Bidirectional Recurrent Neural Network for Depth Video Sequence Based Action Recognition","volume":"32","author":"Liu","year":"2018","journal-title":"Int. J. Pattern Recognit. Artif. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1197","DOI":"10.1007\/s11760-018-1271-3","article-title":"Combining 2D and 3D deep models for action recognition with depth information","volume":"12","author":"Keceli","year":"2018","journal-title":"Signal Image Video Process."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"2639","DOI":"10.1162\/0899766042321814","article-title":"Canonical Correlation Analysis: An Overview with Application to Learning Methods","volume":"16","author":"Hardoon","year":"2004","journal-title":"Neural Comput."},{"key":"ref_23","first-page":"823","article-title":"Cluster Canonical Correlation Analysis","volume":"33","author":"Rasiwasia","year":"2014","journal-title":"JMLR Workshop Conf. Proc."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"188","DOI":"10.1109\/TPAMI.2015.2435740","article-title":"Multi-View Discriminant Analysis","volume":"38","author":"Kan","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Kan, M., Shan, S., and Chen, X. (2016, January 27\u201330). Multi-view Deep Network for Cross-View Classification. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.524"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"531","DOI":"10.1016\/j.imavis.2006.04.014","article-title":"Locality preserving CCA with applications to data visualization and pose estimation","volume":"25","author":"Sun","year":"2007","journal-title":"Image Vis. Comput."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"397","DOI":"10.1016\/j.neucom.2014.06.015","article-title":"A unified multiset canonical correlation analysis framework based on graph embedding for multiple feature extraction","volume":"148","author":"Shen","year":"2015","journal-title":"Neurocomputing"},{"key":"ref_28","unstructured":"Mungoli, N. (2023). Adaptive Feature Fusion: Enhancing Generalization in Deep Learning Models. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"16439","DOI":"10.1007\/s00521-021-06239-5","article-title":"Local-aware spatio-temporal attention network with multi-stage feature fusion for human action recognition","volume":"33","author":"Hou","year":"2021","journal-title":"Neural Comput. Appl."},{"key":"ref_30","unstructured":"Tishby, N., Pereira, F.C., and Bialek, W. (2000). The Information Bottleneck Method. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"663","DOI":"10.1038\/s41592-019-0476-x","article-title":"Markov models-Markov chains","volume":"16","author":"Grewal","year":"2019","journal-title":"Nat. Methods"},{"key":"ref_32","unstructured":"Alemi, A.A., Fischer, I., Dillon, J.V., and Murphy, K. (2016). Deep Variational Information Bottleneck. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Min, Y., Zhang, Y., Chai, X., and Chen, X. (2020, January 13\u201319). An Efficient PointLSTM for Point Clouds Based Gesture Recognition. Proceedings of the 2020 IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00580"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Fan, H., Yang, Y., and Kankanhalli, M. (2021, January 20\u201325). Point 4D Transformer Networks for Spatio-Temporal Modeling in Point Cloud Videos. Proceedings of the 2021 IEEE Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01398"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Li, X., Huang, Q., Zhang, Y., Yang, T., and Wang, Z. (2023). PointMapNet: Point Cloud Feature Map Network for 3D Human Action Recognition. Symmetry, 15.","DOI":"10.3390\/sym15020363"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"8933","DOI":"10.1109\/TII.2022.3223225","article-title":"Real-Time 3-D Human Action Recognition Based on Hyperpoint Sequence","volume":"19","author":"Li","year":"2023","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"872","DOI":"10.1007\/978-3-642-33709-3_62","article-title":"Robust 3D Action Recognition with Random Occupancy Patterns","volume":"7573","author":"Wang","year":"2012","journal-title":"Lect. Notes Comput. Sci."},{"key":"ref_38","unstructured":"Wang, J., Liu, Z., Wu, Y., and Yuan, J. (2012, January 16\u201321). Mining Actionlet Ensemble for Action Recognition with Depth Cameras. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Oreifej, O., and Liu, Z. (2013, January 23\u201328). HON4D: Histogram of Oriented 4D Normals for Activity Recognition from Depth Sequences. Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.98"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Xia, L., and Aggarwal, J.K. (2013, January 23\u201328). Spatio-Temporal Depth Cuboid Similarity Feature for Activity Recognition Using Depth Camera. Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.365"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Tran, Q.D., and Ly, N.Q. (2013, January 10\u201313). Sparse Spatio-Temporal Representation of Joint Shape-Motion Cues for Human Action Recognition in Depth Sequences. Proceedings of the 2013 RIVF International Conference on Computing & Communication Technologies\u2014Research, Innovation, and Vision for Future (RIVF), Hanoi, Vietnam.","DOI":"10.1109\/RIVF.2013.6719903"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"952","DOI":"10.1109\/TCSVT.2014.2302558","article-title":"Body Surface Context: A New Robust Feature for Action Recognition From Depth Videos","volume":"24","author":"Song","year":"2014","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Lu, C., Jia, J., and Tang, C.K. (2014, January 23\u201328). Range-Sample Depth Feature for Action Recognition. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.104"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"1317","DOI":"10.1109\/TMM.2018.2875510","article-title":"Multimodal Learning for Human Action Recognition Via Bimodal\/Multimodal Hybrid Centroid Canonical Correlation Analysis","volume":"21","author":"Elmadany","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"1729881418825093","DOI":"10.1177\/1729881418825093","article-title":"Hierarchical dynamic depth projected difference images-based action recognition in videos with convolutional neural networks","volume":"16","author":"Wu","year":"2019","journal-title":"Int. J. Adv. Robot. Syst."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"1729","DOI":"10.1109\/TCSVT.2018.2855416","article-title":"Dynamic 3D Hand Gesture Recognition by Learning Weighted Depth Motion Maps","volume":"29","author":"Azad","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"4648","DOI":"10.1109\/TIP.2017.2718189","article-title":"Action Recognition Using 3D Histograms of Texture and A Multi-Class Boosting Classifier","volume":"26","author":"Zhang","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1016\/j.neucom.2020.04.034","article-title":"DTMMN: Deep transfer multi -metric network for RGB-D action recognition","volume":"406","author":"Qin","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"135118","DOI":"10.1109\/ACCESS.2020.3006067","article-title":"Depth Sequential Information Entropy Maps and Multi-Label Subspace Learning for Human Action Recognition","volume":"8","author":"Yang","year":"2020","journal-title":"IEEE Access"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"554","DOI":"10.1016\/j.neucom.2014.06.085","article-title":"Multi-perspective and multi-modality joint representation and recognition model for 3D action recognition","volume":"151","author":"Gao","year":"2015","journal-title":"Neurocomputing"},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"1824","DOI":"10.1109\/TCSVT.2017.2655521","article-title":"3D Action Recognition Using Multiscale Energy-Based Global Ternary Image","volume":"28","author":"Liu","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Liu, H., Tian, L., Liu, M., and Tang, H. (2015, January 27\u201330). SDM-BSM: A fusing depth scheme for human action recognition. Proceedings of the 2015 IEEE International Conference on Image Processing (ICIP), Quebec City, QC, Canada.","DOI":"10.1109\/ICIP.2015.7351693"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Xie, S., Sun, C., Huang, J., Tu, Z., and Murphy, K. (2018, January 8\u201314). Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. Proceedings of the European Conference On Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"2010","DOI":"10.1109\/TPAMI.2015.2505311","article-title":"Joint Feature Selection and Subspace Learning for Cross-Modal Retrieval","volume":"38","author":"Wang","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Ryumin, D., Ivanko, D., and Ryumina, E. (2023). Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices. Sensors, 23.","DOI":"10.3390\/s23042284"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Jiang, S., Sun, B., Wang, L., Bai, Y., Li, K., and Fu, Y. (2021, January 19\u201325). Skeleton Aware Multi-modal Sign Language Recognition. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Nashville, TN, USA.","DOI":"10.1109\/CVPRW53098.2021.00380"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Hruz, M., Gruber, I., Kanis, J., Bohacek, M., Hlavac, M., and Krnoul, Z. (2022). One Model is Not Enough: Ensembles for Isolated Sign Language Recognition. Sensors, 22.","DOI":"10.3390\/s22135043"},{"key":"ref_58","unstructured":"Maxim, N., Leonid, V., Ruslan, M., Dmitriy, M., and Iuliia, Z. (2023). Fine-tuning of sign language recognition models: A technical report. arXiv."},{"key":"ref_59","unstructured":"Ryumin, D., Ivanko, D., and Axyonov, A. (2023, January 24\u201326). Cross-Language Transfer Learning Using Visual Information for Automatic Sign Gesture Recognition. Proceedings of the International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, Moscow, Russia."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/15\/12\/2177\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:35:07Z","timestamp":1760132107000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/15\/12\/2177"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,8]]},"references-count":59,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["sym15122177"],"URL":"https:\/\/doi.org\/10.3390\/sym15122177","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,8]]}}}