{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:57:19Z","timestamp":1760241439075,"version":"build-2065373602"},"reference-count":43,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2018,2,20]],"date-time":"2018-02-20T00:00:00Z","timestamp":1519084800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>With the introduction of cost-effective depth sensors, a tremendous amount of research has been devoted to studying human action recognition using 3D motion data. However, most existing methods work in an offline fashion, i.e., they operate on a segmented sequence. There are a few methods specifically designed for online action recognition, which continually predicts action labels as a stream sequence proceeds. In view of this fact, we propose a question: can we draw inspirations and borrow techniques or descriptors from existing offline methods, and then apply these to online action recognition? Note that extending offline techniques or descriptors to online applications is not straightforward, since at least two problems\u2014including real-time performance and sequence segmentation\u2014are usually not considered in offline action recognition. In this paper, we give a positive answer to the question. To develop applicable online action recognition methods, we carefully explore feature extraction, sequence segmentation, computational costs, and classifier selection. The effectiveness of the developed methods is validated on the MSR 3D Online Action dataset and the MSR Daily Activity 3D dataset.<\/jats:p>","DOI":"10.3390\/s18020633","type":"journal-article","created":{"date-parts":[[2018,2,20]],"date-time":"2018-02-20T11:56:13Z","timestamp":1519127773000},"page":"633","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Exploring 3D Human Action Recognition: from Offline to Online"],"prefix":"10.3390","volume":"18","author":[{"given":"Rui","family":"Li","sequence":"first","affiliation":[{"name":"State Key Lab of CAD&amp;CG, Zhejiang University, Hangzhou 310027, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhenyu","family":"Liu","sequence":"additional","affiliation":[{"name":"State Key Lab of CAD&amp;CG, Zhejiang University, Hangzhou 310027, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianrong","family":"Tan","sequence":"additional","affiliation":[{"name":"State Key Lab of CAD&amp;CG, Zhejiang University, Hangzhou 310027, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2018,2,20]]},"reference":[{"key":"ref_1","first-page":"149","article-title":"A survey on human motion analysis from depth data","volume":"Volume 8200","author":"Grzegorzek","year":"2013","journal-title":"Time-of-Flight and Depth Imaging. Sensors, Algorithms, and Applications"},{"key":"ref_2","first-page":"341","article-title":"Vision based human activity recognition: A review","volume":"Volume 513","author":"Angelov","year":"2017","journal-title":"Advances in Computational Intelligence Systems"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1016\/j.imavis.2017.01.010","article-title":"Going deeper into action recognition: A survey","volume":"60","author":"Herath","year":"2017","journal-title":"Image Vision Comput."},{"key":"ref_4","first-page":"1297","article-title":"Real-time human pose recognition in parts from a single depth image","volume":"56","author":"Shotton","year":"2011","journal-title":"IEEE Comput. Vis. Pattern Recognit."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Luo, J., Wang, W., and Qi, H. (2013, January 1\u20138). Group sparsity and geometry constrained dictionary learning for action recognition from depth maps. Proceedings of the IEEE International Conference on Computer Vision, Sydney, NSW, Australia.","DOI":"10.1109\/ICCV.2013.227"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Vemulapalli, R., Arrate, F., and Chellappa, R. (2014, January 23\u201328). Human action recognition by representing 3D human skeletons as points in a Lie group. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.82"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"420","DOI":"10.1007\/s11263-012-0550-7","article-title":"Exploring the trade-off between accuracy and observational latency in action recognition","volume":"101","author":"Ellis","year":"2013","journal-title":"Int. J. Comput. Vis."},{"key":"ref_8","first-page":"1023","article-title":"3-D human action recognition by shape analysis of motion trajectories on Riemannian manifold","volume":"45","author":"Devanne","year":"2014","journal-title":"IEEE Trans. Syst. Man Cybernetics"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TPAMI.2015.2439257","article-title":"Action recognition using rate-invariant analysis of skeletal shape trajectories","volume":"38","author":"Amor","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Yang, X., and Tian, Y. (2012, January 16\u201321). EigenJoints-based action recognition using na\u00efve-bayes-nearest-neighbor. Proceedings of the IEEE Computer Vision and Pattern Recognition Workshops, Providence, RI, USA.","DOI":"10.1109\/CVPRW.2012.6239232"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Zanfir, M., Leordeanu, M., and Sminchisescu, C. (2013, January 1\u20138). The moving pose: An efficient 3D kinematics descriptor for low-latency action recognition and detection. Proceedings of the IEEE International Conference on Computer Vision, Sydney, NSW, Australia.","DOI":"10.1109\/ICCV.2013.342"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Zhu, G., Zhang, L., Shen, P., and Song, J. (2016). An online continuous human action recognition algorithm based on the Kinect sensor. Sensors, 16.","DOI":"10.3390\/s16020161"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1414","DOI":"10.1109\/TPAMI.2013.244","article-title":"Structured time series analysis for human action segmentation and recognition","volume":"36","author":"Gong","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","first-page":"595","article-title":"Continuous gesture recognition from articulated poses","volume":"Volume 8925","author":"Agapito","year":"2014","journal-title":"Computer Vision-ECCV 2014 Workshops. ECCV 2014"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Li, W., Zhang, Z., and Liu, Z. (2010, January 13\u201318). Action recognition based on a bag of 3D points. Proceedings of the IEEE Computer Vision and Pattern Recognition Workshops, San Francisco, CA, USA.","DOI":"10.1109\/CVPRW.2010.5543273"},{"key":"ref_16","unstructured":"Yang, X., Zhang, C., and Tian, Y. (November, January 29). Recognizing actions using depth motion maps-based histograms of oriented gradients. Proceedings of the ACM International Conference on Multimedia, Nara, Japan."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"155","DOI":"10.1007\/s11554-013-0370-1","article-title":"Real-time human action recognition based on depth motion maps","volume":"12","author":"Chen","year":"2013","journal-title":"J. Real-Time Image Process."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Chen, C., Jafari, R., and Kehtarnavaz, N. (2015, January 5\u20139). Action recognition from depth sequences using depth motion maps-based local binary patterns. Proceedings of the IEEE WACV, Waikoloa, HI, USA.","DOI":"10.1109\/WACV.2015.150"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"498","DOI":"10.1109\/THMS.2015.2504550","article-title":"Action recognition from depth maps using deep convolutional neural networks","volume":"46","author":"Wang","year":"2016","journal-title":"IEEE Trans. Human-Machine Syst."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Xia, L., and Aggarwal, J.K. (2013, January 23\u201328). Spatio-temporal depth cuboid similarity feature for activity recognition using depth camera. Proceedings of the IEEE Computer Vision and Pattern Recognition, Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.365"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"5197","DOI":"10.3390\/s150305197","article-title":"Continuous human action recognition using depth-MHI-HOG and a spotter model","volume":"15","author":"Eum","year":"2015","journal-title":"Sensors"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"914","DOI":"10.1109\/TPAMI.2013.198","article-title":"Learning actionlet ensemble for 3D human action recognition","volume":"36","author":"Wang","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ohn-Bar, E., and Trivedi, M.M. (2013, January 23\u201328). Joint angles similarities and HOG2 for action recognition. Proceedings of the IEEE Computer Vision and Pattern Recognition Workshops, Portland, OR, USA.","DOI":"10.1109\/CVPRW.2013.76"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Cremers, D., Reid, I., Saito, H., and Yang, M.H. (2014). Discriminative orderlet mining for real-time recognition of human-object interaction. Computer Vision\u2014ACCV 2014. ACCV 2014, Springer. Lecture Notes in Computer Science.","DOI":"10.1007\/978-3-319-16811-1"},{"key":"ref_25","first-page":"2617","article-title":"Keep it simple and sparse: real-time action recognition","volume":"14","author":"Fanello","year":"2013","journal-title":"J. Mach. Learn. Res."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wu, C., Zhang, J., Savarese, S., and Saxena, A. (2015, January 7). Watch-n-patch: Unsupervised understanding of actions and relations. Proceedings of the IEEE Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299065"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"582","DOI":"10.1109\/TPAMI.2012.137","article-title":"Hierarchical aligned cluster analysis for temporal clustering of human motion","volume":"35","author":"Zhou","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_28","unstructured":"Li, R., Liu, Z., and Tan, J. (2007). Human motion segmentation using collaborative representations of 3D skeletal sequences. IET Comput. Vision."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Fitzgibbon, A., Lazebnik, S., Perona, P., Sato, Y., and Schmid, C. (2012). Kernelized temporal cut for online temporal segmentation and recognition. Computer Vision\u2014ECCV 2012. ECCV 2012. Lecture Notes in Computer Science, Springer.","DOI":"10.1007\/978-3-642-33709-3"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"4311","DOI":"10.1109\/TSP.2006.881199","article-title":"K-SVD: An algorithm for designing of overcomplete dictionaries for sparse representation","volume":"54","author":"Aharon","year":"2006","journal-title":"IEEE Trans. Signal Process."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1111\/j.2517-6161.1996.tb02080.x","article-title":"Regression shrinkage and selection via the lasso","volume":"58","author":"Tibshirani","year":"1996","journal-title":"J. Royal Statistical Soc. B"},{"key":"ref_32","unstructured":"Hoerl, A., and Kennard, R. (1988). Ridge regression. Encyclopedia of Statistical Sciences, Wiley."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"301","DOI":"10.1111\/j.1467-9868.2005.00503.x","article-title":"Regularization and variable selection via the elastic net","volume":"67","author":"Zou","year":"2005","journal-title":"J. Royal Stat. Soc."},{"key":"ref_34","first-page":"1331","article-title":"Human action recognition by learning bases of action attributes and parts","volume":"23","author":"Yao","year":"2011","journal-title":"Int. J. Comput. Vis."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Gong, D., and Medioni, G. (2011, January 6\u201313). Dynamic manifold warping for view invariant action recognition. Proceedings of the IEEE Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126290"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"222","DOI":"10.1016\/j.patcog.2016.07.041","article-title":"Motion segment decomposition of RGB-D sequences for human behavior understanding","volume":"61","author":"Devanne","year":"2017","journal-title":"Pattern Recogn."},{"key":"ref_37","unstructured":"(2017, May 21). Sparse Models, Algorithms and Learning for Large-scale data, SMALLBox. Available online: http:\/\/small-project.eu."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"407","DOI":"10.1214\/009053604000000067","article-title":"Least angle regression","volume":"32","author":"Efron","year":"2004","journal-title":"Ann. Stat."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"1348","DOI":"10.1198\/016214501753382273","article-title":"Variable selection via nonconcave penalized likelihood and its oracle properties","volume":"96","author":"Fan","year":"2001","journal-title":"J. Am. Stat. Assoc."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"971","DOI":"10.1109\/TPAMI.2002.1017623","article-title":"Multiresolution gray-scale and rotation invariant texture classification with local binary patterns","volume":"24","author":"Ojala","year":"2002","journal-title":"IEEE Trans. Pattern Analy. Mach. Intell."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1961189.1961199","article-title":"LIBSVM: A library for support vector machines","volume":"2","author":"Chang","year":"2011","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1016\/j.patcog.2016.05.019","article-title":"RGB-D-based action recognition datasets: A survey","volume":"60","author":"Zhang","year":"2016","journal-title":"Pattern Recogn."},{"key":"ref_43","unstructured":"Zhang, L., Yang, M., and Feng, X. (2011, January 6\u201313). Sparse representation or collaborative representation: Which helps face recognition?. Proceedings of the IEEE International Conference on Computer Vision, Barcelona, Spain."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/2\/633\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T14:55:43Z","timestamp":1760194543000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/2\/633"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,2,20]]},"references-count":43,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2018,2]]}},"alternative-id":["s18020633"],"URL":"https:\/\/doi.org\/10.3390\/s18020633","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2018,2,20]]}}}