{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,29]],"date-time":"2026-07-29T06:39:39Z","timestamp":1785307179200,"version":"3.55.0"},"reference-count":58,"publisher":"MDPI AG","issue":"23","license":[{"start":{"date-parts":[[2021,11,28]],"date-time":"2021-11-28T00:00:00Z","timestamp":1638057600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Human action recognition (HAR) has gained significant attention recently as it can be adopted for a smart surveillance system in Multimedia. However, HAR is a challenging task because of the variety of human actions in daily life. Various solutions based on computer vision (CV) have been proposed in the literature which did not prove to be successful due to large video sequences which need to be processed in surveillance systems. The problem exacerbates in the presence of multi-view cameras. Recently, the development of deep learning (DL)-based systems has shown significant success for HAR even for multi-view camera systems. In this research work, a DL-based design is proposed for HAR. The proposed design consists of multiple steps including feature mapping, feature fusion and feature selection. For the initial feature mapping step, two pre-trained models are considered, such as DenseNet201 and InceptionV3. Later, the extracted deep features are fused using the Serial based Extended (SbE) approach. Later on, the best features are selected using Kurtosis-controlled Weighted KNN. The selected features are classified using several supervised learning algorithms. To show the efficacy of the proposed design, we used several datasets, such as KTH, IXMAS, WVU, and Hollywood. Experimental results showed that the proposed design achieved accuracies of 99.3%, 97.4%, 99.8%, and 99.9%, respectively, on these datasets. Furthermore, the feature selection step performed better in terms of computational time compared with the state-of-the-art.<\/jats:p>","DOI":"10.3390\/s21237941","type":"journal-article","created":{"date-parts":[[2021,12,1]],"date-time":"2021-12-01T01:45:02Z","timestamp":1638323102000},"page":"7941","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":53,"title":["Human Action Recognition: A Paradigm of Best Deep Learning Features Selection and Serial Based Extended Fusion"],"prefix":"10.3390","volume":"21","author":[{"given":"Seemab","family":"Khan","sequence":"first","affiliation":[{"name":"Department of Computer Science, HITEC University Taxila, Txila 47080, Pakistan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6347-4890","authenticated-orcid":false,"given":"Muhammad Attique","family":"Khan","sequence":"additional","affiliation":[{"name":"Department of Computer Science, HITEC University Taxila, Txila 47080, Pakistan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Majed","family":"Alhaisoni","sequence":"additional","affiliation":[{"name":"College of Computer Science and Engineering, University of Ha\u2019il, Ha\u2019il 55211, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7672-1187","authenticated-orcid":false,"given":"Usman","family":"Tariq","sequence":"additional","affiliation":[{"name":"College of Computer Engineering and Science, Prince Sattam Bin Abdulaziz University, Al-Kharaj 11942, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hwan-Seung","family":"Yong","sequence":"additional","affiliation":[{"name":"Department of Computer Science & Engineering, Ewha Womans University, Seoul 120-750, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9062-7493","authenticated-orcid":false,"given":"Ammar","family":"Armghan","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, College of Engineering, Jouf University, Sakakah 72311, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4099-1254","authenticated-orcid":false,"given":"Fayadh","family":"Alenezi","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering, College of Engineering, Jouf University, Sakakah 72311, Saudi Arabia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,11,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Kim, D., Lee, I., Kim, D., and Lee, S. (2021). Action Recognition Using Close-Up of Maximum Activation and ETRI-Activity3D LivingLab Dataset. Sensors, 21.","DOI":"10.3390\/s21206774"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Mishra, O., Kavimandan, P.S., Tripathi, M., Kapoor, R., and Yadav, K. (2021). Human Action Recognition Using a New Hybrid Descriptor. Advances in VLSI, Communication and Signal Processing, Springer.","DOI":"10.1007\/978-981-15-6840-4_43"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"948","DOI":"10.1166\/jmihi.2021.3340","article-title":"Design and Implementation of Human-Computer Interaction Systems Based on Transfer Support Vector Machine and EEG Signal for Depression Patients\u2019 Emotion Recognition","volume":"11","author":"Chen","year":"2021","journal-title":"J. Med. Imaging Health Inform."},{"key":"ref_4","unstructured":"Javed, K., Khan, S.A., Saba, T., Habib, U., Khan, J.A., and Abbasi, A.A. (2020). Human action recognition using fusion of multiview and deep features: An application to video surveillance. Multimed. Tools. Appl., 1\u201327."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Liu, D., Xu, H., Wang, J., Lu, Y., Kong, J., and Qi, M. (2021). Adaptive Attention Memory Graph Convolutional Networks for Skeleton-Based Action Recognition. Sensors, 21.","DOI":"10.3390\/s21206761"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"2217","DOI":"10.32604\/cmc.2021.018103","article-title":"Real-Time Violent Action Recognition Using Key Frames Extraction and Deep Learning","volume":"69","author":"Ahmed","year":"2021","journal-title":"Comput. Mater. Continua"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Wang, J., Cao, D., Wang, J., and Liu, C. (2021). Action Recognition of Lower Limbs Based on Surface Electromyography Weighted Feature Method. Sensors, 21.","DOI":"10.3390\/s21186147"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zin, T.T., Htet, Y., Akagi, Y., Tamura, H., Kondo, K., Araki, S., and Chosa, E. (2021). Real-Time Action Recognition System for Elderly People Using Stereo Depth Camera. Sensors, 21.","DOI":"10.3390\/s21175895"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Farnoosh, A., Wang, Z., Zhu, S., and Ostadabbas, S. (2021). A Bayesian Dynamical Approach for Human Action Recognition. Sensors, 21.","DOI":"10.3390\/s21165613"},{"key":"ref_10","first-page":"237","article-title":"Awareness of voluntary and involuntary causal actions and their outcomes","volume":"2","author":"Buehner","year":"2015","journal-title":"Psychol. Conscious. Theory Res. Pract."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Hassaballah, M., and Hosny, K.M. (2019). Studies in Computational Intelligence. Recent Advances In Computer Vision, Springer.","DOI":"10.1007\/978-3-030-03000-1"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"105986","DOI":"10.1016\/j.asoc.2019.105986","article-title":"Hand-crafted and deep convolutional neural network features fusion and selection strategy: An application to intelligent human action recognition","volume":"87","author":"Sharif","year":"2020","journal-title":"Appl. Soft Comput."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Kolekar, M.H., and Dash, D.P. (2016, January 22\u201325). Hidden markov model based human activity recognition using shape and optical flow based features. Proceedings of the 2016 IEEE Region 10 Conference (TENCON), Singapore.","DOI":"10.1109\/TENCON.2016.7848028"},{"key":"ref_14","unstructured":"Hermansky, H. (December, January 30). TRAP-TANDEM: Data-driven extraction of temporal features from speech. Proceedings of the 2003 IEEE Workshop on Automatic Speech Recognition and Understanding (IEEE Cat. No. 03EX721), St Thomas, VI, USA."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Cabri, J., Pezarat-Correia, P., and Vilas-Boas, J. (2016). The Application of Multiview Human Body Tracking on the Example of Hurdle Clearance. Sport Science Research and Technology Support, Springer.","DOI":"10.1007\/978-3-319-52770-3"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Hassaballah, M., and Awad, A.I. (2020). Deep Learning In Computer Vision: Principles and Applications, CRC Press.","DOI":"10.1201\/9781351003827"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Voulodimos, A., Doulamis, N., Doulamis, A., and Protopapadakis, E. (2018). Deep learning for computer vision: A brief review. Comput. Intell. Neurosci.","DOI":"10.1155\/2018\/7068349"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1109\/MCI.2018.2840738","article-title":"Recent trends in deep learning based natural language processing","volume":"13","author":"Young","year":"2018","journal-title":"IEEE Comput. Intell. Mag."},{"key":"ref_19","unstructured":"Palacio-Ni\u00f1o, J.-O., and Berzal, F. (2019). Evaluation metrics for unsupervised learning algorithms. arXiv."},{"key":"ref_20","first-page":"4061","article-title":"Multi-Layered Deep Learning Features Fusion for Human Action Recognition","volume":"69","author":"Kiran","year":"2021","journal-title":"Comput. Mater. Cont."},{"key":"ref_21","first-page":"3841","article-title":"Video Analytics Framework for Human Action Recognition","volume":"68","author":"Khan","year":"2021","journal-title":"Comput. Mater. Cont."},{"key":"ref_22","first-page":"329","article-title":"Stomach deformities recognition using rank-based deep features selection","volume":"43","author":"Sharif","year":"2019","journal-title":"J. Med. Econ."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Saleem, F., Khan, M.A., Alhaisoni, M., Tariq, U., Armghan, A., Alenezi, F., Choi, J., and Kadry, S. (2021). Human Gait Recognition: A Single Stream Optimal Deep Learning Features Fusion. Sensors, 21.","DOI":"10.3390\/s21227584"},{"key":"ref_24","first-page":"2113","article-title":"Human Gait Recognition Using Deep Learning and Improved Ant Colony Optimization","volume":"70","author":"Khan","year":"2022","journal-title":"Comput. Mater. Cont."},{"key":"ref_25","first-page":"343","article-title":"Human Gait Recognition: A Deep Learning and Best Feature Selection Framework","volume":"70","author":"Mehmood","year":"2022","journal-title":"Comput. Mater. Cont."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.neucom.2020.10.037","article-title":"Skeleton Edge Motion Networks for Human Action Recognition","volume":"423","author":"Wang","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1016\/j.future.2020.10.011","article-title":"Human action identification by a quality-guided fusion of multi-model feature","volume":"116","author":"Bi","year":"2021","journal-title":"Future Gener. Comput. Syst."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"106587","DOI":"10.1016\/j.ymssp.2019.106587","article-title":"Applications of machine learning to machine fault diagnosis: A review and roadmap","volume":"138","author":"Lei","year":"2020","journal-title":"Mech. Syst. Signal Process"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Manivannan, A., Chin, W.C.B., Barrat, A., and Bouffanais, R. (2020). On the challenges and potential of using barometric sensors to track human activity. Sensors, 20.","DOI":"10.3390\/s20236786"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Ahmed Bhuiyan, R., Ahmed, N., Amiruzzaman, M., and Islam, M.R. (2020). A robust feature extraction model for human activity characterization using 3-axis accelerometer and gyroscope data. Sensors, 20.","DOI":"10.3390\/s20236990"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Zhao, B., Li, S., Gao, Y., Li, C., and Li, W. (2020). A Framework of Combining Short-Term Spatial\/Frequency Feature Extraction and Long-Term IndRNN for Activity Recognition. Sensors, 20.","DOI":"10.3390\/s20236984"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"820","DOI":"10.1016\/j.future.2021.06.045","article-title":"Human action recognition using attention based LSTM network with dilated CNN features","volume":"125","author":"Muhammad","year":"2021","journal-title":"Future Gener. Comput. Syst."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Li, C., Xie, C., Zhang, B., Han, J., Zhen, X., and Chen, J. (2021). Memory attention networks for skeleton-based action recognition. IEEE Trans. Neural Netw. Learn. Syst.","DOI":"10.1109\/TNNLS.2021.3061115"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Vedaldi, A., Bischof, H., Brox, T., and Frahm, J.M. (2020). Unsupervised Learning of Optical Flow with Deep Feature Similarity. Computer Vision\u2014ECCV 2020. ECCV 2020, Springer. Lecture Notes in Computer Science.","DOI":"10.1007\/978-3-030-58517-4"},{"key":"ref_35","first-page":"5120","article-title":"$p$-Laplacian regularized sparse coding for human activity recognition","volume":"63","author":"Liu","year":"2016","journal-title":"IEEE Trans. Ind. Electron."},{"key":"ref_36","first-page":"54","article-title":"A Depth Video-based Human Detection and Activity Recognition using Multi-features and Embedded Hidden Markov Models for Health Care Monitoring Systems","volume":"4","author":"Jalal","year":"2017","journal-title":"Int. J. Interact. Multimed. Artif. Intell."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"101224","DOI":"10.1016\/j.ecoinf.2021.101224","article-title":"An evaluation of feature selection methods for environmental data","volume":"61","author":"Effrosynidis","year":"2021","journal-title":"Ecol Inform."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Melhart, D., Liapis, A., and Yannakakis, G.N. (2021). The Affect Game AnnotatIoN (AGAIN) Dataset. arXiv.","DOI":"10.1109\/TAFFC.2022.3188851"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"307","DOI":"10.1016\/j.future.2017.11.029","article-title":"A robust human activity recognition system using smartphone sensors and deep learning","volume":"81","author":"Hassan","year":"2018","journal-title":"Future Gener. Comput. Syst."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"106139","DOI":"10.1016\/j.optlaseng.2020.106139","article-title":"Triple color image encryption based on 2D multiple parameter fractional discrete Fourier transform and 3D Arnold transform","volume":"133","author":"Joshi","year":"2020","journal-title":"Opt. Lasers. Eng."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1016\/j.neucom.2019.10.118","article-title":"A comprehensive survey on support vector machine classification: Applications, challenges and trends","volume":"408","author":"Cervantes","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"17913","DOI":"10.1109\/ACCESS.2018.2817253","article-title":"Human action recognition by learning spatio-temporal features with deep neural networks","volume":"6","author":"Wang","year":"2018","journal-title":"IEEE Access"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"99152","DOI":"10.1109\/ACCESS.2019.2927134","article-title":"A hybrid deep learning model for human activity recognition using multimodal body sensing data","volume":"7","author":"Gumaei","year":"2019","journal-title":"IEEE Access"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"9280","DOI":"10.1109\/JIOT.2019.2911669","article-title":"Adaptive fusion and category-level dictionary learning model for multiview human action recognition","volume":"6","author":"Gao","year":"2019","journal-title":"IEEE Internet Things J."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Khan, M.A., Zhang, Y.-D., Khan, S.A., Attique, M., Rehman, A., and Seo, S. (2020). A resource conscious human action recognition framework using 26-layered deep convolutional neural network. Multimed. Tools. Appl.","DOI":"10.1007\/s11042-020-09408-1"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"56855","DOI":"10.1109\/ACCESS.2020.2982225","article-title":"LSTM-CNN architecture for human activity recognition","volume":"8","author":"Xia","year":"2020","journal-title":"IEEE Access"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"15751","DOI":"10.1007\/s11042-018-7031-0","article-title":"Object detection and classification: A joint selection and fusion strategy of deep convolutional neural network and SIFT point features","volume":"78","author":"Rashid","year":"2019","journal-title":"Multimed. Tools. Appl."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Hussain, N., Sharif, M., Khan, S.A., Albesher, A.A., Saba, T., and Armaghan, A. (2020). A deep neural network and classical features based scheme for objects recognition: An application for machine inspection. Multimed. Tools. Appl., 1\u201323.","DOI":"10.1007\/s11042-020-08852-3"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"58","DOI":"10.1016\/j.patrec.2020.12.015","article-title":"Attributes based skin lesion detection and recognition: A mask RCNN and transfer learning-based deep learning framework","volume":"143","author":"Akram","year":"2021","journal-title":"Pattern Recognit. Lett."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Oquab, M., Bottou, L., Laptev, I., and Sivic, J. (2014, January 23\u201328). Learning and transferring mid-level image representations using convolutional neural networks. Proceedings of the IEEE conference on computer vision and pattern recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.222"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Li, F.-F. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image databas e. Proceedings of the 2009 IEEE conference on computer vision and pattern recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_53","first-page":"1097","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"NIPS"},{"key":"ref_54","first-page":"314","article-title":"Importance of features selection, attributes selection, challenges and future directions for medical imaging data: A review","volume":"125","author":"Naheed","year":"2020","journal-title":"Comput. Sci. Eng."},{"key":"ref_55","first-page":"1","article-title":"Automatic human posture estimation for sport activity recognition with robust body parts detection and entropy markov model","volume":"22","author":"Nadeem","year":"2021","journal-title":"Multimed. Tools. Appl."},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"281","DOI":"10.1007\/s10044-019-00789-0","article-title":"Human action recognition: A framework of statistical weighted segmentation and rank correlation-based selection","volume":"23","author":"Sharif","year":"2020","journal-title":"Pattern Anal. Appl."},{"key":"ref_57","doi-asserted-by":"crossref","first-page":"1377","DOI":"10.1007\/s10044-018-0688-1","article-title":"An implementation of optimized framework for action classification using multilayers neural network on selected fused features","volume":"22","author":"Akram","year":"2019","journal-title":"Pattern Anal. Appl."},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Laptev, I., Marszalek, M., Schmid, C., and Rozenfeld, B. (2008, January 23\u201328). Learning realistic human actions from movies. Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition, Anchorage, AK, USA.","DOI":"10.1109\/CVPR.2008.4587756"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/23\/7941\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:36:57Z","timestamp":1760168217000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/23\/7941"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,11,28]]},"references-count":58,"journal-issue":{"issue":"23","published-online":{"date-parts":[[2021,12]]}},"alternative-id":["s21237941"],"URL":"https:\/\/doi.org\/10.3390\/s21237941","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,11,28]]}}}