{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T08:09:11Z","timestamp":1782893351514,"version":"3.54.5"},"reference-count":43,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2021,10,1]],"date-time":"2021-10-01T00:00:00Z","timestamp":1633046400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,10,1]],"date-time":"2021-10-01T00:00:00Z","timestamp":1633046400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Soft Comput"],"published-print":{"date-parts":[[2022,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Human activity recognition aims to determine actions performed by a human in an image or video. Examples of human activity include standing, running, sitting, sleeping,<jats:italic>etc<\/jats:italic>. These activities may involve intricate motion patterns and undesired events such as falling. This paper proposes a novel deep convolutional long short-term memory (ConvLSTM) network for skeletal-based activity recognition and fall detection. The proposed ConvLSTM network is a sequential fusion of convolutional neural networks (CNNs), long short-term memory (LSTM) networks, and fully connected layers. The acquisition system applies human detection and pose estimation to pre-calculate skeleton coordinates from the image\/video sequence. The ConvLSTM model uses the raw skeleton coordinates along with their characteristic geometrical and kinematic features to construct the novel guided features. The geometrical and kinematic features are built upon raw skeleton coordinates using relative joint position values, differences between joints, spherical joint angles between selected joints, and their angular velocities. The novel spatiotemporal-guided features are obtained using a trained multi-player CNN-LSTM combination. Classification head including fully connected layers is subsequently applied. The proposed model has been evaluated on the KinectHAR dataset having 130,000 samples with 81 attribute values, collected with the help of a Kinect (v2) sensor. Experimental results are compared against the performance of isolated CNNs and LSTM networks. Proposed ConvLSTM have achieved an accuracy of 98.89% that is better than CNNs and LSTMs having an accuracy of 93.89 and 92.75%, respectively. The proposed system has been tested in realtime and is found to be independent of the pose, facing of the camera, individuals, clothing,<jats:italic>etc<\/jats:italic>. The code and dataset will be made publicly available.<\/jats:p>","DOI":"10.1007\/s00500-021-06238-7","type":"journal-article","created":{"date-parts":[[2021,10,1]],"date-time":"2021-10-01T10:17:17Z","timestamp":1633083437000},"page":"877-890","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":81,"title":["Skeleton-based human activity recognition using ConvLSTM and guided feature learning"],"prefix":"10.1007","volume":"26","author":[{"given":"Santosh Kumar","family":"Yadav","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kamlesh","family":"Tiwari","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9128-068X","authenticated-orcid":false,"given":"Hari Mohan","family":"Pandey","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shaik Ali","family":"Akbar","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,10,1]]},"reference":[{"issue":"4","key":"6238_CR1","first-page":"160","volume":"17","author":"B Almaslukh","year":"2017","unstructured":"Almaslukh B, AlMuhtadi J, Artoli A (2017) An effective deep autoencoder approach for online smartphone-based human activity recognition. Int J Comput Sci Netw Secur 17(4):160\u2013165","journal-title":"Int J Comput Sci Netw Secur"},{"issue":"2","key":"6238_CR2","doi-asserted-by":"publisher","first-page":"290","DOI":"10.1109\/TITB.2010.2087385","volume":"15","author":"E Auvinet","year":"2010","unstructured":"Auvinet E, Multon F, Saint-Arnaud A, Rousseau J, Meunier J (2010) Fall detection with multiple cameras: an occlusion-resistant method based on 3-d silhouette vertical distribution. IEEE Trans Inf Technol Biomed 15(2):290\u2013300","journal-title":"IEEE Trans Inf Technol Biomed"},{"issue":"2","key":"6238_CR3","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1109\/72.279181","volume":"5","author":"Y Bengio","year":"1994","unstructured":"Bengio Y, Simard P, Frasconi P (1994) Learning long-term dependencies with gradient descent is difficult. IEEE Trans Neural Netw 5(2):157\u2013166","journal-title":"IEEE Trans Neural Netw"},{"key":"6238_CR4","doi-asserted-by":"publisher","first-page":"148","DOI":"10.1016\/j.patcog.2016.01.020","volume":"55","author":"H Chen","year":"2016","unstructured":"Chen H, Wang G, Xue JH, He L (2016) A novel hierarchical framework for human action recognition. Pattern Recognit 55:148\u2013159","journal-title":"Pattern Recognit"},{"key":"6238_CR5","doi-asserted-by":"crossref","unstructured":"Cippitelli E, Gasparrini S, Gambi E, Spinsante S (2016) A human activity recognition system using skeleton data from rgbd sensors. Comput Intell Neurosci","DOI":"10.1155\/2016\/4351435"},{"key":"6238_CR6","unstructured":"DeSA U, et\u00a0al (2013) World population prospects: the 2012 revision. Population division of the department of economic and social affairs of the United Nations Secretariat, New York 18"},{"key":"6238_CR7","doi-asserted-by":"crossref","unstructured":"Ding Z, Wang P, Ogunbona PO, Li W (2017) Investigation of different skeleton features for cnn-based 3d action recognition. In: 2017 IEEE international conference on multimedia & expo workshops (ICMEW). IEEE, pp 617\u2013622","DOI":"10.1109\/ICMEW.2017.8026286"},{"key":"6238_CR8","unstructured":"Du Y, Wang W, Wang L (2015) Hierarchical recurrent neural network for skeleton based action recognition. In: proceedings of the IEEE conference on computer vision and pattern recognition, pp 1110\u20131118"},{"issue":"2","key":"6238_CR9","doi-asserted-by":"publisher","first-page":"2756","DOI":"10.3390\/s140202756","volume":"14","author":"S Gasparrini","year":"2014","unstructured":"Gasparrini S, Cippitelli E, Spinsante S, Gambi E (2014) A depth-based fall detection system using a kinect\u00ae sensor. Sensors 14(2):2756\u20132775","journal-title":"Sensors"},{"key":"6238_CR10","doi-asserted-by":"crossref","unstructured":"Grushin A, Monner DD, Reggia JA, Mishra A (2013) Robust human action recognition via long short-term memory. In: the 2013 international joint conference on neural networks (IJCNN). IEEE, pp 1\u20138","DOI":"10.1109\/IJCNN.2013.6706797"},{"key":"6238_CR11","unstructured":"Hammerla NY, Halloran S, Pl\u00f6tz T (2016) Deep, convolutional, and recurrent models for human activity recognition using wearables. arXiv preprint arXiv:1604.08880"},{"key":"6238_CR12","doi-asserted-by":"crossref","unstructured":"Hou L, Wan W, Han K, Muhammad R, Yang M (2016) Human detection and tracking over camera networks: a review. In: 2016 international conference on audio, language and image processing (ICALIP). IEEE, pp 574\u2013580","DOI":"10.1109\/ICALIP.2016.7846643"},{"issue":"2","key":"6238_CR13","doi-asserted-by":"publisher","first-page":"173","DOI":"10.1007\/s10015-017-0422-x","volume":"23","author":"M Inoue","year":"2018","unstructured":"Inoue M, Inoue S, Nishida T (2018) Deep recurrent neural network for mobile human activity recognition with high throughput. Artif Life Robot 23(2):173\u2013185","journal-title":"Artif Life Robot"},{"issue":"6","key":"6238_CR14","doi-asserted-by":"publisher","first-page":"731","DOI":"10.1109\/LSP.2017.2690339","volume":"24","author":"Q Ke","year":"2017","unstructured":"Ke Q, An S, Bennamoun M, Sohel F, Boussaid F (2017) Skeletonnet: mining deep part features for 3-d action recognition. IEEE Signal Process Lett 24(6):731\u2013735","journal-title":"IEEE Signal Process Lett"},{"issue":"7553","key":"6238_CR15","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y LeCun","year":"2015","unstructured":"LeCun Y, Bengio Y, Hinton G (2015) Deep learning. Nature 521(7553):436\u2013444","journal-title":"Nature"},{"key":"6238_CR16","doi-asserted-by":"crossref","unstructured":"Lee I, Kim D, Kang S, Lee S (2017) Ensemble deep learning for skeleton-based action recognition using temporal sliding lstm networks. In: proceedings of the IEEE international conference on computer vision, pp 1012\u20131020","DOI":"10.1109\/ICCV.2017.115"},{"key":"6238_CR17","unstructured":"Li C, Wang P, Wang S, Hou Y, Li W (2017) Skeleton-based action recognition using lstm and cnn. In: 2017 IEEE international conference on multimedia & expo workshops (ICMEW). IEEE, pp 585\u2013590"},{"key":"6238_CR18","doi-asserted-by":"crossref","unstructured":"Li J, Luong MT, Jurafsky D (2015) A hierarchical neural autoencoder for paragraphs and documents. arXiv preprint arXiv:1506.01057","DOI":"10.3115\/v1\/P15-1107"},{"issue":"4","key":"6238_CR19","doi-asserted-by":"publisher","first-page":"2579","DOI":"10.1007\/s11071-019-05149-5","volume":"97","author":"MW Li","year":"2019","unstructured":"Li MW, Geng J, Hong WC, Zhang LD (2019) Periodogram estimation based on lssvr-ccpso compensation for forecasting ship motion. Nonlinear Dyn 97(4):2579\u20132594","journal-title":"Nonlinear Dyn"},{"key":"6238_CR20","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1007\/978-3-319-13817-6_11","volume-title":"Mining intelligence and knowledge exploration","author":"Y Li","year":"2014","unstructured":"Li Y, Shi D, Ding B, Liu D (2014) Unsupervised feature learning for human activity recognition using smartphone sensors. Mining intelligence and knowledge exploration. Springer, Berlin, pp 99\u2013107"},{"issue":"12","key":"6238_CR21","doi-asserted-by":"publisher","first-page":"3007","DOI":"10.1109\/TPAMI.2017.2771306","volume":"40","author":"J Liu","year":"2017","unstructured":"Liu J, Shahroudy A, Xu D, Kot AC, Wang G (2017) Skeleton-based action recognition using spatio-temporal lstm network with trust gates. IEEE Trans Pattern Anal Mach Intell 40(12):3007\u20133021","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"1","key":"6238_CR22","doi-asserted-by":"publisher","first-page":"115","DOI":"10.3390\/s16010115","volume":"16","author":"FJ Ord\u00f3\u00f1ez","year":"2016","unstructured":"Ord\u00f3\u00f1ez FJ, Roggen D (2016) Deep convolutional and lstm recurrent neural networks for multimodal wearable activity recognition. Sensors 16(1):115","journal-title":"Sensors"},{"key":"6238_CR23","unstructured":"O\u2019Shea K, Nash R (2015) An introduction to convolutional neural networks. arXiv preprint arXiv:1511.08458"},{"issue":"6","key":"6238_CR24","doi-asserted-by":"publisher","first-page":"976","DOI":"10.1016\/j.imavis.2009.11.014","volume":"28","author":"R Poppe","year":"2010","unstructured":"Poppe R (2010) A survey on vision-based human action recognition. Image Vis Comput 28(6):976\u2013990","journal-title":"Image Vis Comput"},{"issue":"5","key":"6238_CR25","doi-asserted-by":"publisher","first-page":"611","DOI":"10.1109\/TCSVT.2011.2129370","volume":"21","author":"C Rougier","year":"2011","unstructured":"Rougier C, Meunier J, St-Arnaud A, Rousseau J (2011) Robust video surveillance for fall detection based on human shape deformation. IEEE Trans Circuits Syst Video Technol 21(5):611\u2013622","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"6238_CR26","doi-asserted-by":"crossref","unstructured":"Singh MS, Pondenkandath V, Zhou B, Lukowicz P, Liwickit M (2017) Transforming sensor data to the image domain for deep learning-an application to footstep detection. In: 2017 international joint conference on neural networks (IJCNN). IEEE, pp 2665\u20132672","DOI":"10.1109\/IJCNN.2017.7966182"},{"issue":"5","key":"6238_CR27","doi-asserted-by":"publisher","first-page":"290","DOI":"10.1136\/ip.2005.011015","volume":"12","author":"JA Stevens","year":"2006","unstructured":"Stevens JA, Corso PS, Finkelstein EA, Miller TR (2006) The costs of fatal and non-fatal falls among older adults. Injury Prevent 12(5):290\u2013295","journal-title":"Injury Prevent"},{"key":"6238_CR28","unstructured":"Uden L, P\u00e9rez J, Herrera F, Rodr\u0131guez J (2018) Advances in intelligent systems and computing: Preface. Advances in Intelligent Systems and Computing 172"},{"key":"6238_CR29","doi-asserted-by":"crossref","unstructured":"Vepakomma P, De D, Das SK, Bhansali S (2015) A-wristocracy: Deep learning on wrist-worn sensing for recognition of user complex activities. In: 2015 IEEE 12th international conference on wearable and implantable body sensor networks (BSN). IEEE, pp 1\u20136","DOI":"10.1109\/BSN.2015.7299406"},{"key":"6238_CR30","doi-asserted-by":"crossref","unstructured":"Wang A, Chen G, Shang C, Zhang M, Liu L (2016) Human activity recognition in a smart home environment with stacked denoising autoencoders. In: international conference on web-age information management. Springer, pp 29\u201340","DOI":"10.1007\/978-3-319-47121-1_3"},{"key":"6238_CR31","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1016\/j.patrec.2018.02.010","volume":"119","author":"J Wang","year":"2019","unstructured":"Wang J, Chen Y, Hao S, Peng X, Hu L (2019) Deep learning for sensor-based activity recognition: a survey. Pattern Recognit Lett 119:3\u201311","journal-title":"Pattern Recognit Lett"},{"issue":"5","key":"6238_CR32","doi-asserted-by":"publisher","first-page":"914","DOI":"10.1109\/TPAMI.2013.198","volume":"36","author":"J Wang","year":"2013","unstructured":"Wang J, Liu Z, Wu Y, Yuan J (2013) Learning actionlet ensemble for 3d human action recognition. IEEE Trans Pattern Anal Mach Intell 36(5):914\u2013927","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"4","key":"6238_CR33","doi-asserted-by":"publisher","first-page":"498","DOI":"10.1109\/THMS.2015.2504550","volume":"46","author":"P Wang","year":"2015","unstructured":"Wang P, Li W, Gao Z, Zhang J, Tang C, Ogunbona PO (2015) Action recognition from depth maps using deep convolutional neural networks. IEEE Trans Human-Mach Syst 46(4):498\u2013509","journal-title":"IEEE Trans Human-Mach Syst"},{"issue":"2","key":"6238_CR34","doi-asserted-by":"publisher","first-page":"224","DOI":"10.1016\/j.cviu.2010.10.002","volume":"115","author":"D Weinland","year":"2011","unstructured":"Weinland D, Ronfard R, Boyer E (2011) A survey of vision-based methods for action representation, segmentation and recognition. Comput Vis Image Understand 115(2):224\u2013241","journal-title":"Comput Vis Image Understand"},{"issue":"4","key":"6238_CR35","doi-asserted-by":"publisher","first-page":"1077","DOI":"10.1109\/TCSVT.2018.2818151","volume":"29","author":"J Weng","year":"2018","unstructured":"Weng J, Weng C, Yuan J, Liu Z (2018) Discriminative spatio-temporal pattern discovery for 3d action recognition. IEEE Trans Circuits Syst Video Technol 29(4):1077\u20131089","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"6238_CR36","doi-asserted-by":"publisher","first-page":"106970","DOI":"10.1016\/j.knosys.2021.106970","volume":"223","author":"SK Yadav","year":"2021","unstructured":"Yadav SK, Tiwari K, Pandey HM, Akbar SA (2021) A review of multimodal human activity recognition with special emphasis on classification, applications, challenges and future directions. Knowl-Based Syst 223:106970","journal-title":"Knowl-Based Syst"},{"key":"6238_CR37","doi-asserted-by":"crossref","unstructured":"Yao S, Hu S, Zhao Y, Zhang A, Abdelzaher T (2017) Deepsense: A unified deep learning framework for time-series mobile sensing data processing. In: proceedings of the 26th international conference on world wide web, pp 351\u2013360","DOI":"10.1145\/3038912.3052577"},{"key":"6238_CR38","doi-asserted-by":"crossref","unstructured":"Yuan J, Liu Z, Wu Y (2009) Discriminative subvolume search for efficient action detection. In: 2009 IEEE conference on computer vision and pattern recognition. IEEE, pp 2442\u20132449","DOI":"10.1109\/CVPR.2009.5206671"},{"key":"6238_CR39","doi-asserted-by":"publisher","first-page":"86","DOI":"10.1016\/j.patcog.2016.05.019","volume":"60","author":"J Zhang","year":"2016","unstructured":"Zhang J, Li W, Ogunbona PO, Wang P, Tang C (2016) Rgb-d-based action recognition datasets: a survey. Pattern Recognit 60:86\u2013105","journal-title":"Pattern Recognit"},{"key":"6238_CR40","doi-asserted-by":"crossref","unstructured":"Zhang P, Lan C, Xing J, Zeng W, Xue J, Zheng N (2017) View adaptive recurrent neural networks for high performance human action recognition from skeleton data. In: proceedings of the IEEE international conference on computer vision, pp 2117\u20132126","DOI":"10.1109\/ICCV.2017.233"},{"key":"6238_CR41","doi-asserted-by":"crossref","unstructured":"Zhang P, Xue J, Lan C, Zeng W, Gao Z, Zheng N (2018) Adding attentiveness to the neurons in recurrent neural networks. In: proceedings of the European conference on computer vision (ECCV), pp 135\u2013151","DOI":"10.1007\/978-3-030-01240-3_9"},{"key":"6238_CR42","doi-asserted-by":"publisher","first-page":"3047","DOI":"10.1109\/TNNLS.2019.2935173","volume":"31","author":"X Zhang","year":"2019","unstructured":"Zhang X, Xu C, Tian X, Tao D (2019) Graph edge convolutional neural networks for skeleton-based action recognition. IEEE Trans Neural Netw Learn Syst 31:3047\u20133060","journal-title":"IEEE Trans Neural Netw Learn Syst"},{"key":"6238_CR43","doi-asserted-by":"crossref","unstructured":"Zhu W, Lan C, Xing J, Zeng W, Li Y, Shen L, Xie X (2016) Co-occurrence feature learning for skeleton based action recognition using regularized deep lstm networks. arXiv preprint arXiv:1603.07772","DOI":"10.1609\/aaai.v30i1.10451"}],"container-title":["Soft Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00500-021-06238-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00500-021-06238-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00500-021-06238-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,10]],"date-time":"2023-01-10T22:57:19Z","timestamp":1673391439000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00500-021-06238-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,1]]},"references-count":43,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2022,1]]}},"alternative-id":["6238"],"URL":"https:\/\/doi.org\/10.1007\/s00500-021-06238-7","relation":{},"ISSN":["1432-7643","1433-7479"],"issn-type":[{"value":"1432-7643","type":"print"},{"value":"1433-7479","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,10,1]]},"assertion":[{"value":"1 September 2021","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 October 2021","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that there is no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}