{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T07:22:25Z","timestamp":1776064945312,"version":"3.50.1"},"reference-count":43,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2019,4,2]],"date-time":"2019-04-02T00:00:00Z","timestamp":1554163200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"This work was supported by Institute for Information and communications Technology Promotion (IITP) grant funded by the Korea government (MSIT) (No. 2016-0-00406, SIAT CCTV Cloud Platform).","award":["No. 2016-0-00406, SIAT CCTV Cloud Platform"],"award-info":[{"award-number":["No. 2016-0-00406, SIAT CCTV Cloud Platform"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Human action recognition plays a significant part in the research community due to its emerging applications. A variety of approaches have been proposed to resolve this problem, however, several issues still need to be addressed. In action recognition, effectively extracting and aggregating the spatial-temporal information plays a vital role to describe a video. In this research, we propose a novel approach to recognize human actions by considering both deep spatial features and handcrafted spatiotemporal features. Firstly, we extract the deep spatial features by employing a state-of-the-art deep convolutional network, namely Inception-Resnet-v2. Secondly, we introduce a novel handcrafted feature descriptor, namely Weber\u2019s law based Volume Local Gradient Ternary Pattern (WVLGTP), which brings out the spatiotemporal features. It also considers the shape information by using gradient operation. Furthermore, Weber\u2019s law based threshold value and the ternary pattern based on an adaptive local threshold is presented to effectively handle the noisy center pixel value. Besides, a multi-resolution approach for WVLGTP based on an averaging scheme is also presented. Afterward, both these extracted features are concatenated and feed to the Support Vector Machine to perform the classification. Lastly, the extensive experimental analysis shows that our proposed method outperforms state-of-the-art approaches in terms of accuracy.<\/jats:p>","DOI":"10.3390\/s19071599","type":"journal-article","created":{"date-parts":[[2019,4,3]],"date-time":"2019-04-03T03:39:28Z","timestamp":1554262768000},"page":"1599","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":29,"title":["Feature Fusion of Deep Spatial Features and Handcrafted Spatiotemporal Features for Human Action Recognition"],"prefix":"10.3390","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7718-5627","authenticated-orcid":false,"given":"Md Azher","family":"Uddin","sequence":"first","affiliation":[{"name":"Department of Computer Science and Engineering, Kyung Hee University, Global Campus, Yongin 17104, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Young-Koo","family":"Lee","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Kyung Hee University, Global Campus, Yongin 17104, Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,4,2]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Baumann, F., Liao, J., Ehlers, A., and Rosenhahn, B. (2014, January 26\u201329). Computation strategies for volume local binary patterns applied to action recognition. Proceedings of the 11th IEEE International Conference on Advanced Video and Signal-Based Surveillance (AVSS), Seoul, Korea.","DOI":"10.1109\/AVSS.2014.6918646"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"54","DOI":"10.1016\/j.neucom.2015.03.097","article-title":"Recognizing human actions using novel space-time volume binary patterns","volume":"173","author":"Baumann","year":"2016","journal-title":"Neurocomputing"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"5125","DOI":"10.1016\/j.eswa.2010.09.137","article-title":"Local Ternary Patterns from Three Orthogonal Planes for human action classification","volume":"38","author":"Laptev","year":"2011","journal-title":"Expert Syst. Appl."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1016\/j.eswa.2017.01.008","article-title":"Realistic action recognition with salient foreground trajectories","volume":"75","author":"Yi","year":"2017","journal-title":"Expert Syst. Appl."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"915","DOI":"10.1109\/TPAMI.2007.1110","article-title":"Dynamic Texture Recognition Using Local Binary Patterns with an Application to Facial Expressions","volume":"29","author":"Zhao","year":"2007","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"971","DOI":"10.1109\/TPAMI.2002.1017623","article-title":"Multiresolution gray-scale and rotation invariant texture classification with local binary patterns","volume":"7","author":"Ojala","year":"2002","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"21157","DOI":"10.1109\/ACCESS.2017.2759225","article-title":"Human Action Recognition Using Adaptive Local Motion Descriptor in Spark","volume":"5","author":"Uddin","year":"2017","journal-title":"IEEE Access"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lan, T., Zhu, Y., Zamir, A.R., and Savarese, S. (2016, January 7\u201313). Action recognition by hierarchical mid-level action elements. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.517"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"219","DOI":"10.1007\/s11263-015-0846-5","article-title":"Action recognition with improved trajectories","volume":"119","author":"Wang","year":"2016","journal-title":"Int. J. Comput. Vis."},{"key":"ref_10","unstructured":"Simonyan, K., and Zisserman, A. (2014, January 8\u201313). Two-Stream Convolutional Networks for Action Recognition in Videos. Proceedings of the 27th International Conference on Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Wang, H., Klaser, A., Schmid, C., and Liu, C.-L. (2011, January 20\u201325). Action recognition by dense trajectories. Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition, Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995407"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., and Li, F.-F. (2014, January 23\u201328). Large-scale Video Classification with Convolutional Neural Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.223"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Yue-Hei Ng, J., Hausknecht, M., Vijayanarasimhan, S., Vinyals, O., Monga, R., and Toderici, G. (2015, January 7\u201312). Beyond Short Snippets: Deep Networks for Video Classification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299101"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Mattivi, R., and Shao, L. (2009, January 2\u20134). Human Action Recognition Using LBP-TOP as Sparse Spatio-Temporal Feature Descriptor. Proceedings of the 13th International Conference on Computer Analysis of Images and Patterns, M\u00fcnster, Germany.","DOI":"10.1007\/978-3-642-03767-2_90"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A.A. (2017, January 4\u20139). Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning. Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1023\/A:1009715923555","article-title":"A tutorial on support vector machines for pattern recognition","volume":"2","author":"Burges","year":"1998","journal-title":"Data Min. Knowl. Discov."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Sch\u00fcldt, C., Laptev, I., and Caputo, B. (2004, January 23\u201326). Recognizing Human Actions: A Local SVM Approach. Proceedings of the 17th International Conference on Pattern Recognition, Cambridge, UK.","DOI":"10.1109\/ICPR.2004.1334462"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Rodriguez, M.D., Ahmed, J., and Shah, M. (2008, January 23\u201328). Action MACH: A Spatio-temporal Maximum Average Correlation Height Filter for Action Recognition. Proceedings of the Computer Vision and Pattern Recognition, Anchorage, AK, USA.","DOI":"10.1109\/CVPR.2008.4587727"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Soomro, K., and Zamir, A.R. (2014). Action Recognition in Realistic Sports Videos. Computer Vision in Sports, Springer International Publishing.","DOI":"10.1007\/978-3-319-09396-3_9"},{"key":"ref_20","unstructured":"Ryoo, M.S., and Aggarwal, J.K. (October, January 29). Spatio-Temporal Relationship Match: Video Structure Comparison for Recognition of Complex Human Activities. Proceedings of the 12th International Conference on Computer Vision, Kyoto, Japan."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Marszalek, M., Laptev, I., and Schmid, C. (2009, January 20\u201325). Actions in context. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPRW.2009.5206557"},{"key":"ref_22","unstructured":"Soomro, K., Zamir, A.R., and Shah, M. (arXiv, 2012). UCF101: A Dataset of 101 Human Action Classes From Videos in The Wild, arXiv."},{"key":"ref_23","unstructured":"Yeffet, L., and Wolf, L. (October, January 29). Local Trinary Patterns for human action recognition. Proceedings of the 12th International Conference on Computer Vision, Kyoto, Japan."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1388","DOI":"10.1587\/transinf.2017EDL8006","article-title":"A Novel 3D Gradient LBP Descriptor for Action Recognition","volume":"100","author":"Guo","year":"2017","journal-title":"IEICE Trans. Inf. Syst."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"800","DOI":"10.1109\/TCSVT.2018.2816960","article-title":"ML-HDP: A Hierarchical Bayesian Nonparametric Model for Recognizing Human Actions in Video","volume":"29","author":"Tu","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Dalal, N., Triggs, B., and Schmid, C. (2006, January 7\u201313). Human detection using oriented histograms of flow and appearance. Proceedings of the 9th European conference on Computer Vision (ECCV), Graz, Austria.","DOI":"10.1007\/11744047_33"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1007\/s11263-012-0594-8","article-title":"Dense trajectories and motion boundary descriptors for action recognition","volume":"103","author":"Wang","year":"2013","journal-title":"Int. J. Comput. Vis."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Chakraborty, B., Holte, M.B., Moeslun, T.B., Gonz\u00e0lez, J., and Xavier Roca, F. (2011, January 6\u201313). A selective spatio-temporal interest point detector for human action recognition in complex scenes. Proceedings of the International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126443"},{"key":"ref_29","unstructured":"Chen, M., and Hauptmann, A. (2009). MoSIFT: Recognizing Human actions in Surveillance Videos. [Ph.D. Dissertation, Carnegie Mellon Universtiy]."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Ohnishi, K., Hidaka, M., and Harada, T. (2016, January 15\u201319). Improved Dense Trajectory with Cross Streams. Proceedings of the 24th ACM International Conference on Multimedia, Amsterdam, The Netherlands.","DOI":"10.1145\/2964284.2967222"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Wang, L., Qiao, Y., and Tang, X. (2015, January 7\u201312). Action Recognition with Trajectory-Pooled Deep-Convolutional Descriptors. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299059"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"507","DOI":"10.1007\/s11042-017-5251-3","article-title":"Action recognition with multi-scale trajectory-pooled 3D convolutional descriptors","volume":"78","author":"Lu","year":"2017","journal-title":"Multimed. Tools Appl."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Yao, G., Lei, T., Zhong, J., and Jiang, P. (2018). Learning multi-temporal-scale deep information for action recognition. Appl. Intell., 1\u201313.","DOI":"10.1007\/s10489-018-1347-3"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wang, L., Zang, J., Zhang, Q., Niu, Z., Hua, G., and Zheng, N. (2018). Action Recognition by an Attention-Aware Temporal Weighted Convolutional Neural Network. Sensors, 18.","DOI":"10.3390\/s18071979"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Girdhar, R., Ramanan, D., Gupta, A., Sivic, J., and Russell, B. (2017, January 21\u201326). ActionVLAD: Learning spatio-temporal aggregation for action classification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.337"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"4933","DOI":"10.1109\/TIP.2018.2846664","article-title":"Sequential Video VLAD: Training the Aggregation Locally and Temporally","volume":"27","author":"Xu","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"1839","DOI":"10.1109\/TCSVT.2017.2682196","article-title":"Pooling the Convolutional Layers in Deep ConvNets for Video Action Recognition","volume":"28","author":"Zhao","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 27\u201330). Rethinking the Inception Architecture for Computer Vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_40","unstructured":"Jain, A.K. (1989). Fundamentals of Digital Signal Processing, Prentice-Hall."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"1705","DOI":"10.1109\/TPAMI.2009.155","article-title":"WLD: A Robust Local Image Descriptor","volume":"32","author":"Chen","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_42","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20136). ImageNet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_43","unstructured":"Simonyan, K., and Zisserman, A. (2015, January 7\u20139). Very deep convolutional networks for large-scale image recognition. Proceedings of the International Conference on Learning Representations, San Diego, CA, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/7\/1599\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:42:25Z","timestamp":1760186545000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/7\/1599"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,4,2]]},"references-count":43,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2019,4]]}},"alternative-id":["s19071599"],"URL":"https:\/\/doi.org\/10.3390\/s19071599","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,4,2]]}}}