{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T04:10:46Z","timestamp":1773202246134,"version":"3.50.1"},"reference-count":39,"publisher":"MDPI AG","issue":"17","license":[{"start":{"date-parts":[[2020,8,19]],"date-time":"2020-08-19T00:00:00Z","timestamp":1597795200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100014220","name":"NSFC-Henan Joint Fund","doi-asserted-by":"publisher","award":["U1804152"],"award-info":[{"award-number":["U1804152"]}],"id":[{"id":"10.13039\/501100014220","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>To achieve the satisfactory performance of human action recognition, a central task is to address the sub-action sharing problem, especially in similar action classes. Nevertheless, most existing convolutional neural network (CNN)-based action recognition algorithms uniformly divide video into frames and then randomly select the frames as inputs, ignoring the distinct characteristics among different frames. In recent years, depth videos have been increasingly used for action recognition, but most methods merely focus on the spatial information of the different actions without utilizing temporal information. In order to address these issues, a novel energy-guided temporal segmentation method is proposed here, and a multimodal fusion strategy is employed with the proposed segmentation method to construct an energy-guided temporal segmentation network (EGTSN). Specifically, the EGTSN had two parts: energy-guided video segmentation and a multimodal fusion heterogeneous CNN. The proposed solution was evaluated on a public large-scale NTU RGB+D dataset. Comparisons with state-of-the-art methods demonstrate the effectiveness of the proposed network.<\/jats:p>","DOI":"10.3390\/s20174673","type":"journal-article","created":{"date-parts":[[2020,8,19]],"date-time":"2020-08-19T09:22:31Z","timestamp":1597828951000},"page":"4673","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Energy-Guided Temporal Segmentation Network for Multimodal Human Action Recognition"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3705-3822","authenticated-orcid":false,"given":"Qiang","family":"Liu","sequence":"first","affiliation":[{"name":"School of Information Engineering, Zhengzhou University, Zhengzhou 450000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1261-1282","authenticated-orcid":false,"given":"Enqing","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Information Engineering, Zhengzhou University, Zhengzhou 450000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lei","family":"Gao","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, Ryerson University, Toronto, ON M5B 2K3, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6291-5212","authenticated-orcid":false,"given":"Chengwu","family":"Liang","sequence":"additional","affiliation":[{"name":"School of Information Engineering, Zhengzhou University, Zhengzhou 450000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hao","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Information Engineering, Zhengzhou University, Zhengzhou 450000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,8,19]]},"reference":[{"key":"ref_1","unstructured":"Niu, W., Long, J., Han, D., and Wang, Y. (2004, January 27\u201330). Human activity detection and recognition for video surveillance. Proceedings of the IEEE International Conference on Multimedia and Expo (ICME), Taipei, Taiwan."},{"key":"ref_2","unstructured":"Pickering, C.A., Burnham, K.J., and Richardson, M.J. (2007, January 28\u201329). A Research study of hand gesture recognition technologies and applications for human vehicle interaction. Proceedings of the Institution of Engineering and Technology Conference on Automotive Electronics, Warwick, UK."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"262","DOI":"10.14445\/22315381\/IJETT-V19P245","article-title":"Hand gesture recognition for real time human machine interaction system","volume":"19","author":"Poonam","year":"2015","journal-title":"Int. J. Eng. Trends Technol."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Weiyao, L., Mingting, S., Radha, P., and Zhengyou, Z. (2008, January 18\u201321). Human activity recognition for video surveillance. Proceedings of the 2008 IEEE International Symposium on Circuits and Systems, Seattle, WA, USA.","DOI":"10.1109\/ISCAS.2008.4542023"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1806","DOI":"10.1109\/TSMC.2018.2850149","article-title":"Deep convolutional neural networks for human action recognition using depth maps and postures","volume":"49","author":"Kamel","year":"2019","journal-title":"IEEE Trans. Syst. Man Cybern. Syst."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1839","DOI":"10.1109\/TCSVT.2017.2682196","article-title":"Pooling the convolutional layers in deep convnets for video action recognition","volume":"28","author":"Zhao","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_7","unstructured":"Jamie, S., Andrew, F., Mat, C., Toby, S., Mark, F., Richard, M., Alex, K., and Andrew, B. (2011, January 20\u201325). Real-time human pose recognition in parts from single depth images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1016\/j.cviu.2017.11.008","article-title":"DAAL: Deep activation-based attribute learning for action recognition in depth videos","volume":"167","author":"Zhang","year":"2018","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1051","DOI":"10.1109\/TMM.2018.2818329","article-title":"depth pooling based large-scale 3-D action recognition with convolutional neural networks","volume":"20","author":"Wang","year":"2018","journal-title":"IEEE Trans. Multimed."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Sahoo, S.P., and Ari, S. (2019, January 17\u201320). Depth estimated history image based appearance representation for human action recognition. Proceedings of the TENCON 2019\u20142019 IEEE Region 10 Conference (TENCON), Kochi, India.","DOI":"10.1109\/TENCON.2019.8929687"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Kong, Y., and Fu, Y. (2015, January 8\u201310). Bilinear heterogeneous information machine for RGB-D action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298708"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1016\/j.cviu.2015.10.001","article-title":"Learning hierarchical 3D kernel descriptors for RGB-D action recognition","volume":"144","author":"Kong","year":"2016","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_13","first-page":"14","article-title":"Learning principal orientations and residual descriptor for action recognition","volume":"86","author":"Lei","year":"2017","journal-title":"Pattern Recognit."},{"key":"ref_14","unstructured":"Simonyan, K., and Zisserman, A. (2014). Two-stream convolutional networks for action recognition in videos. Advances in Neural Information Processing Systems 27 (NIPS 2014), Neural Information Processing Systems Foundation."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., and Van Gool, L. (2016, January 8\u201316). Temporal segment networks: Towards good practices for deep action recognition. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherland.","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 7\u201313). Learning spatiotemporal features with 3D convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Shou, Z., Lin, X., Kalantidis, Y., Sevilla-Lara, L., Rohrbach, M., Chang, S.-F., and Yan, Z. (2019, January 15\u201320). DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00136"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zolfaghari, M., Singh, K., and Brox, T. (2018, January 8\u201314). ECO: Efficient convolutional network for online video understanding. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01216-8_43"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1423","DOI":"10.1109\/TCSVT.2018.2830102","article-title":"Semantic cues enhanced multimodality multistream CNN for action recognition","volume":"29","author":"Tu","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zolfaghari, M., Oliveira, G.L., Sedaghat, N., and Brox, T. (2017, January 22\u201329). Chained multi-stream networks exploiting pose, motion, and appearance for action classification and detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.316"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Jian, S. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the ieee conference on computer vision and pattern recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R., Doll\u00e1r, P., Tu, Z., and He, K. (2017, January 21\u201326). Aggregated residual transformations for deep neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.634"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wang, H., and Schmid, C. (2013, January 25\u201327). Action Recognition with Improved Trajectories. Proceedings of the IEEE International Conference on Computer Vision, Portland, OR, USA.","DOI":"10.1109\/ICCV.2013.441"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"2186","DOI":"10.1109\/TPAMI.2016.2640292","article-title":"Jointly learning heterogeneous features for RGB-D activity recognition","volume":"39","author":"Hu","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Laptev, I., Marsza\u0142ek, M., Schmid, C., and Rozenfeld, B. (2008, January 23\u201328). Learning realistic human actions from movies. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Anchorage, AK, USA.","DOI":"10.1109\/CVPR.2008.4587756"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"3007","DOI":"10.1109\/TPAMI.2017.2771306","article-title":"Skeleton-based action recognition using spatio-temporal LSTM network with trust gates","volume":"40","author":"Liu","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"677","DOI":"10.1109\/TPAMI.2016.2599174","article-title":"Long-term recurrent convolutional networks for visual recognition and description","volume":"39","author":"Donahue","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wu, Z., Wang, X., Jiang, Y., Ye, H., and Xue, X. (2015, January 26\u201330). Modeling spatial-temporal clues in a hybrid deep learning framework for video classification. Proceedings of the 23rd ACM International Conference on Multimedia, Brisbane, Australia.","DOI":"10.1145\/2733373.2806222"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Shahroudy, A., Liu, J., Ng, T.-T., and Wang, G. (2016, January 27\u201330). NTU RGB+D: A large scale dataset for 3D human activity Analysis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.115"},{"key":"ref_30","unstructured":"Zach, C., Pock, T., and Bischof, H. (2007, January 12\u201314). A Duality based approach for realtime TV-L1 optical flow. Proceedings of the Joint Pattern Recognition Symposium, Heidelberg, Germany."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Lin, T., Zhao, X., Su, H., Wang, C., and Yang, M. (2018, January 8\u201314). BSN: Boundary sensitive network for temporal action proposal generation. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01225-0_1"},{"key":"ref_32","unstructured":"Yu, W., Yang, K., Bai, Y., Xiao, T., Yao, H., and Rui, Y. (2016, January 19\u201324). Visualizing and comparing AlexNet and VGG using deconvolutional layers. Proceedings of the International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_33","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 6\u201311). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the 32nd international Conference on Machine Learning, Lille, France."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Li, F.-F. (2009, January 20\u201325). ImageNet: A large-scale hierarchical image database. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wang, P., Li, Z., Hou, Y., and Li, W. (2016, January 15\u201319). Action Recognition Based on Joint Trajectory Maps Using Convolutional Neural Networks. Proceedings of the 24th ACM International Conference on Multimedia, Amsterdam, The Netherlands.","DOI":"10.1145\/2964284.2967191"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Han, Y., Chung, S.-L., Chen, S.-F., and Su, S.F. (2018, January 7\u201310). Two-Stream LSTM for action recognition with RGB-D-Based hand-crafted features and feature combination. Proceedings of the 2018 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Miyazaki, Japan.","DOI":"10.1109\/SMC.2018.00600"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"148658","DOI":"10.1109\/ACCESS.2019.2945632","article-title":"Robust multi-feature learning for skeleton-based action recognition","volume":"7","author":"Wang","year":"2019","journal-title":"IEEE Access"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"97757","DOI":"10.1109\/ACCESS.2020.2996779","article-title":"Multi-stream and enhanced spatial-temporal graph convolution network for skeleton-based action recognition","volume":"8","author":"Li","year":"2020","journal-title":"IEEE Access"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"1633","DOI":"10.1109\/LSP.2019.2942739","article-title":"Action machine: Toward person-centric action recognition in videos","volume":"26","author":"Zhu","year":"2019","journal-title":"IEEE Signal Process. Lett."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/17\/4673\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T10:03:03Z","timestamp":1760176983000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/17\/4673"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,8,19]]},"references-count":39,"journal-issue":{"issue":"17","published-online":{"date-parts":[[2020,9]]}},"alternative-id":["s20174673"],"URL":"https:\/\/doi.org\/10.3390\/s20174673","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,8,19]]}}}