{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,11,8]],"date-time":"2024-11-08T05:23:33Z","timestamp":1731043413394,"version":"3.28.0"},"reference-count":44,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2017,11,7]],"date-time":"2017-11-07T00:00:00Z","timestamp":1510012800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2017,11,7]],"date-time":"2017-11-07T00:00:00Z","timestamp":1510012800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Image Video Proc."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In this paper, we address the problem of classifying activities of daily living (ADL) in video. The basic idea of the proposed method is to treat each human activity in the video as a temporal sequence of points on a Riemannian manifold and classify such time series with a geodesic-based kernel. The main novelties of this paper are summarized as follows: (a) for each frame of a video, low-level features of body pose and human-object interaction are unified by a covariance matrix, i.e., a manifold point in the space of symmetric positive definite (SPD) matrices <jats:inline-formula><jats:alternatives><jats:tex-math>$Sym_{+}^{d}$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                      <mml:mtext>Sy<\/mml:mtext>\n                      <mml:msubsup>\n                        <mml:mrow>\n                          <mml:mi>m<\/mml:mi>\n                        <\/mml:mrow>\n                        <mml:mrow>\n                          <mml:mo>+<\/mml:mo>\n                        <\/mml:mrow>\n                        <mml:mrow>\n                          <mml:mi>d<\/mml:mi>\n                        <\/mml:mrow>\n                      <\/mml:msubsup>\n                    <\/mml:math><\/jats:alternatives><\/jats:inline-formula>; (b) a time-dependent bag-of-words (BoW+T) model is built, where its codebook is generated by clustering per-frame covariance matrices on <jats:inline-formula><jats:alternatives><jats:tex-math>$Sym_{+}^{d}$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                      <mml:mtext>Sy<\/mml:mtext>\n                      <mml:msubsup>\n                        <mml:mrow>\n                          <mml:mi>m<\/mml:mi>\n                        <\/mml:mrow>\n                        <mml:mrow>\n                          <mml:mo>+<\/mml:mo>\n                        <\/mml:mrow>\n                        <mml:mrow>\n                          <mml:mi>d<\/mml:mi>\n                        <\/mml:mrow>\n                      <\/mml:msubsup>\n                    <\/mml:math><\/jats:alternatives><\/jats:inline-formula>; (c) for each video, high-level BoW+T features are extracted from its corresponding sequence of per-frame covariance matrices; and (d) for activity classification, a positive definite kernel is formulated, taking into account the underlying geometry of our BoW+T features, i.e., the unit <jats:italic>n<\/jats:italic>-sphere. Experiments were conducted on two video datasets. The first dataset contains 8 activity classes with a total of 943 videos, and the second one contains 7 activity classes with a total of 224 videos. The proposed method achieved high accuracy (average 89.66%) and small false alarms (average 1.43%) on the first dataset. Comparison with six exisiting methods on the second dataset showed further evidence on the effectiveness of the proposed method.<\/jats:p>","DOI":"10.1186\/s13640-017-0220-3","type":"journal-article","created":{"date-parts":[[2024,11,7]],"date-time":"2024-11-07T15:10:59Z","timestamp":1730992259000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Time-dependent bag of words on manifolds for geodesic-based classification of video activities towards assisted living and healthcare"],"prefix":"10.1186","volume":"2017","author":[{"given":"Yixiao","family":"Yun","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Irene Yu-Hua","family":"Gu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2017,11,7]]},"reference":[{"issue":"5","key":"220_CR1","doi-asserted-by":"publisher","first-page":"826","DOI":"10.1109\/TCSVT.2013.2280849","volume":"24","author":"Lin","year":"2014","unstructured":"Lin, et al., A new network-based algorithm for human activity recognition in videos. IEEE Trans. Circ. Syst. Video Technol. (T-CSVT).24(5), 826\u2013841 (2014).","journal-title":"IEEE Trans. Circ. Syst. Video Technol. (T-CSVT)."},{"issue":"4","key":"220_CR2","doi-asserted-by":"publisher","first-page":"1569","DOI":"10.1109\/TIP.2014.2302677","volume":"23","author":"I Everts","year":"2014","unstructured":"I Everts, T JC van Gemert, Gevers, Evaluation of color spatio-temporal interest points for human action recognition. IEEE Trans. Image Process. (T-IP).23(4), 1569\u20131580 (2014).","journal-title":"IEEE Trans. Image Process. (T-IP)."},{"issue":"2\/3","key":"220_CR3","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1007\/s11263-005-1838-7","volume":"64","author":"I Laptev","year":"2005","unstructured":"I Laptev, On space-time interest points. Int. J. Comput. Vis. (IJCV). 64(2\/3), 107\u2013123 (2005).","journal-title":"Int. J. Comput. Vis. (IJCV)"},{"issue":"12","key":"220_CR4","doi-asserted-by":"publisher","first-page":"2344","DOI":"10.1109\/LSP.2015.2480097","volume":"22","author":"G Zhang","year":"2015","unstructured":"G Zhang, M Piccardi, Structural SVM with partial ranking for activity segmentation and classification. IEEE Signal Process. Lett. 22(12), 2344\u20132348 (2015).","journal-title":"IEEE Signal Process. Lett"},{"issue":"4","key":"220_CR5","doi-asserted-by":"publisher","first-page":"800","DOI":"10.1109\/TPAMI.2015.2465955","volume":"38","author":"MR Amer","year":"2016","unstructured":"MR Amer, S Todorovic, Sum product networks for activity recognition. IEEE Trans. Pattern. Anal. Mach. Intell. (T-PAMI).38(4), 800\u2013813 (2016).","journal-title":"IEEE Trans. Pattern. Anal. Mach. Intell. (T-PAMI)."},{"issue":"3","key":"220_CR6","doi-asserted-by":"publisher","first-page":"541","DOI":"10.1109\/TCSVT.2014.2376139","volume":"26","author":"H Zhang","year":"2016","unstructured":"H Zhang, LE Parker, CoDe4D: color-depth local spatio-temporal features for human activity recognition from RGB-D videos. IEEE Trans. Circ. Syst. Video Technol. 26(3), 541\u2013555 (2016).","journal-title":"IEEE Trans. Circ. Syst. Video Technol"},{"key":"220_CR7","doi-asserted-by":"crossref","unstructured":"A Karpathy, et al, in Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). large-scale video classification with convolutional neural networks, (2014), pp. 1725\u20131732.","DOI":"10.1109\/CVPR.2014.223"},{"key":"220_CR8","doi-asserted-by":"crossref","unstructured":"J Donahue, et al, in Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). long-term recurrent convolutional networks for visual recognition and description, (2015), pp. 2625\u20132634.","DOI":"10.1109\/CVPR.2015.7298878"},{"key":"220_CR9","doi-asserted-by":"crossref","unstructured":"M Baccouche, et al, in Proceedings of International Workshop on Human Behavior Understanding (HBU). sequential deep learning for human activity recognition, (2011), pp. 29\u201339.","DOI":"10.1007\/978-3-642-25446-8_4"},{"key":"220_CR10","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4419-9982-5","volume-title":"Introduction to smooth manifolds","author":"JM Lee","year":"2012","unstructured":"JM Lee, Introduction to smooth manifolds (Springer, New York, 2012)."},{"issue":"1","key":"220_CR11","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1007\/s11263-005-3222-z","volume":"66","author":"X Pennec","year":"2006","unstructured":"X Pennec, P Fillard, N Ayache, A Riemannian framework for tensor computing. Int. J. Comput. Vis. (IJCV).66(1), 41\u201366 (2006).","journal-title":"Int. J. Comput. Vis. (IJCV)."},{"issue":"1","key":"220_CR12","doi-asserted-by":"publisher","first-page":"328","DOI":"10.1137\/050637996","volume":"29","author":"V Arsigny","year":"2008","unstructured":"V Arsigny, P Fillard, et al., Geometric means in a novel vector space structure on symmetric-positive definite matrices. SIAM. J. Matrix Anal. Appl. (SJMAEL).29(1), 328\u2013347 (2008).","journal-title":"J. Matrix Anal. Appl. (SJMAEL)."},{"issue":"1","key":"220_CR13","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s11263-008-0195-8","volume":"84","author":"R Subbarao","year":"2009","unstructured":"R Subbarao, P Meer, Nonlinear mean shift over Riemannian manifolds. Int. J. Comput. Vis. (IJCV).84(1), 1\u201320 (2009).","journal-title":"Int. J. Comput. Vis. (IJCV)."},{"issue":"18","key":"220_CR14","doi-asserted-by":"publisher","first-page":"1420","DOI":"10.1049\/el.2015.1345","volume":"51","author":"A Traumann","year":"2015","unstructured":"A Traumann, et al., Accurate 3D measurement using optical depth information. Electron. Lett. 51(18), 1420\u20131422 (2015).","journal-title":"Electron. Lett"},{"key":"220_CR15","doi-asserted-by":"crossref","unstructured":"O Tuzel, F Porikli, P Meer, in Proceedings of European Conference on Computer Vision (ECCV). Region covariance: a fast descriptor for detection and classification, (2006), pp. 589\u2013600.","DOI":"10.1007\/11744047_45"},{"issue":"10","key":"220_CR16","doi-asserted-by":"publisher","first-page":"1713","DOI":"10.1109\/TPAMI.2008.75","volume":"30","author":"O Tuzel","year":"2008","unstructured":"O Tuzel, F Porikli, P Meer, Pedestrian detection via classification on Riemannian manifolds. IEEE Trans. Pattern. Anal. Mach. Intell. (T-PAMI).30(10), 1713\u20131727 (2008).","journal-title":"IEEE Trans. Pattern. Anal. Mach. Intell. (T-PAMI)."},{"key":"220_CR17","doi-asserted-by":"crossref","DOI":"10.1201\/b11847","volume-title":"Differential geometry of manifolds","author":"ST Lovett","year":"2010","unstructured":"ST Lovett, Differential geometry of manifolds, 1st edition (A K Peters\/CRC Press, Natick, 2010)."},{"key":"220_CR18","doi-asserted-by":"crossref","unstructured":"S Jayasumana, et al, in Proceedings of IEEE International Conference on, Digital Image Computing: Techniques and Applications (DICTA). Combining multiple manifold-valued descriptors for improved object recognition, (2013), pp. 1\u20136.","DOI":"10.1109\/DICTA.2013.6691493"},{"issue":"2","key":"220_CR19","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","volume":"60","author":"DG Lowe","year":"2004","unstructured":"DG Lowe, Distinctive image features from scale-invariant keypoints. Int. J. Comput. Vis. (IJCV).60(2), 91\u2013110 (2004).","journal-title":"Int. J. Comput. Vis. (IJCV)."},{"key":"220_CR20","first-page":"886","volume":"1","author":"N Dadal","year":"2005","unstructured":"N Dadal, B Triggs, Histograms of oriented gradients for human detection. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR).1:, 886\u2013893 (2005).","journal-title":"IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR)."},{"issue":"7","key":"220_CR21","doi-asserted-by":"publisher","first-page":"971","DOI":"10.1109\/TPAMI.2002.1017623","volume":"24","author":"T Ojala","year":"2002","unstructured":"T Ojala, M Pietik\u00e4inen, T M\u00e4enp\u00e4\u00e4, Multiresolution gray-scale and rotation invariant texture classification with local binary patterns. IEEE Trans. Pattern. Anal Mach. Intell. (T-PAMI).24(7), 971\u2013987 (2002).","journal-title":"IEEE Trans. Pattern. Anal Mach. Intell. (T-PAMI)."},{"key":"220_CR22","unstructured":"L Fei-Fei, P Perona, in Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). A Bayesian hierarchical model for learning natural scene categories, (2005)."},{"issue":"9","key":"220_CR23","doi-asserted-by":"publisher","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","volume":"32","author":"PF Felzenszwalb","year":"2010","unstructured":"PF Felzenszwalb, et al., Object detection with discriminatively trained part based models. IEEE Trans. Pattern. Anal. Mach. Intell. (T-PAMI).32(9), 1627\u20131645 (2010).","journal-title":"IEEE Trans. Pattern. Anal. Mach. Intell. (T-PAMI)."},{"key":"220_CR24","unstructured":"S Lazebnik, C Schmid, J Ponce, in Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Beyond bags of features: spatial pyramid matching for recognizing natural scene categories, (2006)."},{"key":"220_CR25","doi-asserted-by":"crossref","unstructured":"F Perronnin, J S\u00e1nchez, T Mensink, in Proceedings of European Conference on Computer Vision (ECCV). Improving the fisher kernel for large-scale image classification, (2010).","DOI":"10.1007\/978-3-642-15561-1_11"},{"key":"220_CR26","doi-asserted-by":"crossref","unstructured":"H J\u00e9gou, M Douze, C Schmid, in Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Aggregating local descriptors into a compact image representation, (2010).","DOI":"10.1109\/CVPR.2010.5540039"},{"key":"220_CR27","doi-asserted-by":"crossref","unstructured":"S Gudmundsson, TP Runarsson, S Sigurdsson, in Proceedings of IEEE International Joint Conference on Neural Networks (IJCNN). Support vector machines and dynamic time warping for time series, (2008).","DOI":"10.1109\/IJCNN.2008.4634188"},{"key":"220_CR28","doi-asserted-by":"crossref","unstructured":"L Chen, R Ng, in Proceedings of International Conference on Very Large Data Bases (VLDB). On the marriage of Lp-norms and edit distance, (2004), pp. 792\u2013803.","DOI":"10.1016\/B978-012088469-8\/50070-X"},{"issue":"2","key":"220_CR29","doi-asserted-by":"publisher","first-page":"306","DOI":"10.1109\/TPAMI.2008.76","volume":"31","author":"P-F Marteau","year":"2008","unstructured":"P-F Marteau, Time warp edit distance with stiffness adjustment for time series matching. IEEE Trans. Pattern. Anal. Mach. Intell. (T-PAMI).31(2), 306\u2013318 (2008).","journal-title":"IEEE Trans. Pattern. Anal. Mach. Intell. (T-PAMI)."},{"issue":"6","key":"220_CR30","doi-asserted-by":"publisher","first-page":"1121","DOI":"10.1109\/TNNLS.2014.2333876","volume":"26","author":"P-F Marteau","year":"2015","unstructured":"P-F Marteau, S Gibet, On recursive edit distance kernels with application to time series classification. IEEE Trans. Neural Netw. Learn. Syst. (T-NNLS).26(6), 1121\u20131133 (2015).","journal-title":"IEEE Trans. Neural Netw. Learn. Syst. (T-NNLS)."},{"issue":"3","key":"220_CR31","doi-asserted-by":"publisher","first-page":"1171","DOI":"10.1214\/009053607000000677","volume":"36","author":"T Hofmann","year":"2008","unstructured":"T Hofmann, B Sch\u00f6lkopf, AJ Smola, Kernel methods in machine learning. Ann. Stat. 36(3), 1171\u20131220 (2008).","journal-title":"Ann. Stat"},{"key":"220_CR32","unstructured":"M Cuturi, in Proceedings of International Conference on Machine Learning (ICML). Fast global alignment kernel, (2011)."},{"key":"220_CR33","doi-asserted-by":"crossref","unstructured":"MA Livingston, et al, in Proceedings of IEEE Virtual Reality Workshop (VRW). Performance measurements for the Microsoft Kinect skeleton, (2012), pp. 119\u2013120.","DOI":"10.1109\/VR.2012.6180911"},{"key":"220_CR34","unstructured":"X Chen, A Yuille, in Proceedings of Advances in Neural Information Processing Systems (NIPS). Articulated pose estimation by a graphical model with image dependent pairwise relations, (2014)."},{"key":"220_CR35","doi-asserted-by":"crossref","unstructured":"W Shen, in Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Object skeleton extraction in natural images by fusing scale-associated deep side outputs, (2016).","DOI":"10.1109\/CVPR.2016.31"},{"key":"220_CR36","doi-asserted-by":"crossref","unstructured":"G Yu, Z Liu, J Yuan, in Asian Conference on Computer Vision (ACCV). Discriminative orderlet mining for real-time recognition of human-object interaction, (2014).","DOI":"10.1007\/978-3-319-16814-2_4"},{"issue":"3","key":"220_CR37","doi-asserted-by":"publisher","first-page":"27:1","DOI":"10.1145\/1961189.1961199","volume":"2","author":"CC Chang","year":"2011","unstructured":"CC Chang, CJ Lin, LIBSVM: a library for support vector machines. ACM Trans. Intell. Syst. Technol.2(3), 27:1\u201327:27 (2011).","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"220_CR38","unstructured":"J Wang, et al., in IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Mining actionlet ensemble for action recognition with depth cameras, (2012)."},{"key":"220_CR39","doi-asserted-by":"crossref","unstructured":"L Xia, JK Aggarwal, in Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Spatio-temporal depth cuboid similarity feature for activity recognition using depth camera, (2013).","DOI":"10.1109\/CVPR.2013.365"},{"key":"220_CR40","doi-asserted-by":"crossref","unstructured":"X Yang, in IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). EigenJoints-based action recognition using Naive-Bayes-Nearest-Neighbor, (2012).","DOI":"10.1109\/CVPRW.2012.6239232"},{"key":"220_CR41","doi-asserted-by":"crossref","unstructured":"M Zanfir, M Leordeanu, C Sminchisescu, in Proceedings of IEEE International Conference on Computer Vision (ICCV). The moving pose: an efficient 3d kinematics descriptor for low-latency action recognition and detection, (2013).","DOI":"10.1109\/ICCV.2013.342"},{"issue":"3","key":"220_CR42","doi-asserted-by":"publisher","first-page":"522","DOI":"10.1090\/S0002-9947-1938-1501980-0","volume":"44","author":"IJ Schoenberg","year":"1938","unstructured":"IJ Schoenberg, Metric spaces, positivedefinitefunctions. Trans. Am. Math. Soc. (T-AMS).44(3), 522\u2013536 (1938).","journal-title":"Math. Soc. (T-AMS)."},{"key":"220_CR43","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-1128-0","volume-title":"Ressel, Harmonic analysis on semigroups","author":"JPR C Berg","year":"1984","unstructured":"JPR C Berg, P Christensen, Ressel, Harmonic analysis on semigroups (Springer, New York, 1984)."},{"issue":"12","key":"220_CR44","doi-asserted-by":"publisher","first-page":"2406","DOI":"10.1080\/03081087.2015.1015401","volume":"63","author":"P J\u00f3ziak","year":"2015","unstructured":"P J\u00f3ziak, Conditionally strictly negative definite kernels. Linear and Multilinear Algebra. 63(12), 2406\u20132418 (2015).","journal-title":"Linear and Multilinear Algebra"}],"container-title":["EURASIP Journal on Image and Video Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13640-017-0220-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13640-017-0220-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13640-017-0220-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,7]],"date-time":"2024-11-07T15:22:06Z","timestamp":1730992926000},"score":1,"resource":{"primary":{"URL":"https:\/\/jivp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13640-017-0220-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,11,7]]},"references-count":44,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2017,12]]}},"alternative-id":["220"],"URL":"https:\/\/doi.org\/10.1186\/s13640-017-0220-3","relation":{},"ISSN":["1687-5281"],"issn-type":[{"value":"1687-5281","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,11,7]]},"assertion":[{"value":"13 May 2017","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 October 2017","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 November 2017","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Not applicable.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}},{"value":"Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Publisher\u2019s Note"}}],"article-number":"72"}}