{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T17:14:19Z","timestamp":1740158059912,"version":"3.37.3"},"reference-count":45,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2021,10,27]],"date-time":"2021-10-27T00:00:00Z","timestamp":1635292800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,10,27]],"date-time":"2021-10-27T00:00:00Z","timestamp":1635292800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62002053"],"award-info":[{"award-number":["62002053"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003453","name":"Natural Science Foundation of Guangdong Province","doi-asserted-by":"publisher","award":["2021A1515011866"],"award-info":[{"award-number":["2021A1515011866"]}],"id":[{"id":"10.13039\/501100003453","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007471","name":"Applied Basic Research Foundation of Yunnan Province","doi-asserted-by":"publisher","award":["2020A1515110504"],"award-info":[{"award-number":["2020A1515110504"]}],"id":[{"id":"10.13039\/100007471","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Guangdong Basic and Applied Basic Research Projects","award":["2019A1515111082"],"award-info":[{"award-number":["2019A1515111082"]}]},{"name":"Social welfare major project of Zhongshan","award":["2019B2010"],"award-info":[{"award-number":["2019B2010"]}]},{"name":"Social Welfare Major Project of Zhongshan","award":["420S36"],"award-info":[{"award-number":["420S36"]}]},{"name":"Social welfare major project of Zhongshan","award":["2019B2011"],"award-info":[{"award-number":["2019B2011"]}]},{"name":"Fund for high level talents afforded by University of Electronic Science and Technology of China, Zhongshan Institute","award":["419YKQN15","417YKQ12"],"award-info":[{"award-number":["419YKQN15","417YKQ12"]}]},{"name":"Achievement cultivation project of Zhongshan Industrial Technology Research Institute","award":["419N26"],"award-info":[{"award-number":["419N26"]}]},{"name":"the Science and Technology Foundation of Guangdong Province","award":["2021A0101180005"],"award-info":[{"award-number":["2021A0101180005"]}]},{"name":"Young Innovative Talents Project of Education Department of Guangdong Province","award":["2019KQNCX186"],"award-info":[{"award-number":["2019KQNCX186"]}]},{"name":"Young innovative talents project of Education Department of Guangdong Province","award":["2018KQNCX337"],"award-info":[{"award-number":["2018KQNCX337"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Ambient Intell Human Comput"],"published-print":{"date-parts":[[2023,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Based on dynamic mode decomposition (DMD), a new empirical feature for quasi-few-shot setting (QFSS) skeleton-based action recognition (SAR) is proposed in this study. DMD linearizes the system and extracts the modes in the form of flattened system matrix or stacked eigenvalues, named the DMD feature. The DMD feature has three advantages. The first advantage is its translational and rotational invariance with respect to the change in the localization and pose of the camera. The second one is its clear physical meaning, that is, if a skeleton trajectory was treated as the output of a nonlinear closed-loop system, then the modes of the system represent the intrinsic dynamic property of the motion. Finally, the last one is its compact length and its simple calculation without training. The information contained by the DMD feature is not as complete as that of the feature extracted using a deep convolutional neural network (CNN). However, the DMD feature can be concatenated with CNN features to greatly improve their performance in QFSS tasks, in which we do not have adequate samples to train a deep CNN directly or numerous support sets for standard few-shot learning methods. Four QFSS datasets of SAR named CMU, Badminton, miniNTU-xsub, and miniNTU-xview, are established based on the widely used public datasets to validate the performance of the DMD feature. A group of experiments is conducted to analyze intrinsic properties of DMD, whereas another group focuses on its auxiliary functions. Experimental results show that the DMD feature can improve the performance of most typical CNN features in QFSS SAR tasks.<\/jats:p>","DOI":"10.1007\/s12652-021-03567-1","type":"journal-article","created":{"date-parts":[[2021,10,27]],"date-time":"2021-10-27T04:03:02Z","timestamp":1635307382000},"page":"7159-7172","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Action recognition based on dynamic mode decomposition"],"prefix":"10.1007","volume":"14","author":[{"given":"Shuai","family":"Dong","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weixi","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6481-5458","authenticated-orcid":false,"given":"Kun","family":"Zou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,10,27]]},"reference":[{"key":"3567_CR1","doi-asserted-by":"crossref","unstructured":"Cao Z, Sheikh T, Shih-En S, Yaser W (2017) Realtime multi-person 2D pose estimation using part affinity fields. In: IEEE conference on computer vision and pattern recognition, pp 7291\u20137299","DOI":"10.1109\/CVPR.2017.143"},{"key":"3567_CR2","unstructured":"CMU (2013) CMU graphics lab motion capture database"},{"key":"3567_CR3","unstructured":"Diba A, Fayyaz M, Sharma V, Karami AH, Arzani MM, Yousefzadeh R, Gool LV (2017) Temporal 3D ConvNets: new architecture and transfer learning for video classification. arXiv: 171108200 pp. 1\u201310"},{"key":"3567_CR4","doi-asserted-by":"publisher","unstructured":"Feichtenhofer C, Pinz A, Zisserman A (2016) Convolutional two-stream network fusion for video action recognition. IEEE computer society conference on computer vision and pattern recognition. pp. 1933\u20131941. https:\/\/doi.org\/10.1109\/CVPR.2016.213","DOI":"10.1109\/CVPR.2016.213"},{"key":"3567_CR5","doi-asserted-by":"crossref","unstructured":"Feichtenhofer C, Fan H, Malik J, He K (2018) Slowfast networks for video recognition. In: IEEE\/CVF international conference on computer vision, pp. 6201\u20136210","DOI":"10.1109\/ICCV.2019.00630"},{"key":"3567_CR6","first-page":"37","volume-title":"Long short-term memory","author":"A Graves","year":"2012","unstructured":"Graves A (2012) Long short-term memory. Springer, Berlin, pp 37\u201345"},{"key":"3567_CR7","doi-asserted-by":"crossref","unstructured":"Guo M, Chou E, Huang DA, Song S, Yeung S, Fei-Fei L (2018) Neural graph matching networks for few-shot 3D action recognition. European conference on computer vision. Munich, Germany, pp. 673\u2013689","DOI":"10.1007\/978-3-030-01246-5_40"},{"key":"3567_CR8","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. IEEE conference on computer vision and pattern recognition. Las Vegas, USA, pp 771\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"3567_CR9","doi-asserted-by":"publisher","first-page":"28","DOI":"10.1016\/j.inffus.2021.01.008","volume":"71","author":"A Holzinger","year":"2021","unstructured":"Holzinger A, Malle B, Saranti A, Pfeifer B (2021) Towards multi-modal causability with Graph Neural Networks enabling information fusion for explainable AI. Inf Fusion 71:28\u201337. https:\/\/doi.org\/10.1016\/j.inffus.2021.01.008","journal-title":"Inf Fusion"},{"issue":"3","key":"3567_CR10","doi-asserted-by":"publisher","first-page":"807","DOI":"10.1109\/TCSVT.2016.2628339","volume":"28","author":"Y Hou","year":"2018","unstructured":"Hou Y, Li Z, Wang P, Li W (2018) Skeleton optical spectra-based action recognition using convolutional neural networks. IEEE Trans Circuits Syst Video Technol 28(3):807\u2013811. https:\/\/doi.org\/10.1109\/TCSVT.2016.2628339","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"3567_CR11","unstructured":"Jasani B, Mazagonwalla A (2019) Skeleton based zero shot action recognition in joint pose-language semantic space. arXiv: 191111344 pp. 1\u20138, arXiv: 1911.11344v1"},{"key":"3567_CR12","unstructured":"Kay W, Carreira J, Simonyan K, Zhang B, Hillier C, Vijayanarasimhan S, Viola F, Green T, Back T, Natsev P, Suleyman M, Zisserman A (2017) The Kinetics human action video dataset. arXiv: 170506950. pp. 1\u201322"},{"key":"3567_CR13","doi-asserted-by":"crossref","unstructured":"Kim TS, Reiter A (2017) Interpretable 3D human action analysis with temporal convolutional networks. IEEE conference on computer vision and pattern recognition workshops, pp. 1623\u20131631","DOI":"10.1109\/CVPRW.2017.207"},{"key":"3567_CR14","unstructured":"Kong Y, Fu Y (2018) Human action recognition and prediction: a survey. arXiv: 180611230 13(9):1\u201319"},{"key":"3567_CR15","unstructured":"Li B, He M, Cheng X, Chen Y, Dai Y (2017a) Skeleton based action recognition using translation-scale invariant image mapping and multi-scale deep CNN. In: IEEE international conference on multimedia and expo workshops, pp. 601\u2013604"},{"key":"3567_CR16","unstructured":"Li C, Zhong Q, Xie D, Pu S (2017b) Skeleton-based action recognition with convolutional neural networks. IEEE international conference on multimedia and expo workshops. China, Hong Kong, pp. 597\u2013600"},{"key":"3567_CR17","doi-asserted-by":"crossref","unstructured":"Li L, Zheng W, Zhang Z, Huang Y, Wang L (2019) Relational network for skeleton-based action recognition. IEEE international conference on multimedia and expo, pp .826\u2013831. arXiv: 1805.02556v1","DOI":"10.1109\/ICME.2019.00147"},{"key":"3567_CR18","doi-asserted-by":"publisher","unstructured":"Lin J, Gan C, Han S (2019) TSM: temporal shift module for efficient video understanding. In: IEEE\/CVF international conference on computer vision (ICCV), pp. 7082\u20137092. https:\/\/doi.org\/10.1109\/ICCV.2019.00718","DOI":"10.1109\/ICCV.2019.00718"},{"key":"3567_CR19","doi-asserted-by":"publisher","unstructured":"Lin J, Gan C, Wang K, Han S (2020) TSM: Temporal shift module for efficient and scalable video understanding on edge devices. IEEE transactions on pattern analysis and machine intelligence, p. 1, https:\/\/doi.org\/10.1109\/TPAMI.2020.3029799","DOI":"10.1109\/TPAMI.2020.3029799"},{"key":"3567_CR20","doi-asserted-by":"crossref","unstructured":"Liu J, Wang G, Hu P, Duan Ly, Kot AC (2017) Global context-aware attention LSTM networks for 3D action recognition. IEEE conference on computer vision and pattern recognition. pp, 1647\u20131656","DOI":"10.1109\/CVPR.2017.391"},{"key":"3567_CR21","doi-asserted-by":"publisher","unstructured":"Liu R, Shen J, Wang H, Chen C, Cheung SC, Asari V (2020) Attention mechanism exploits temporal contexts: real-time 3D human pose reconstruction. Proceedings of the IEEE computer society conference on computer vision and pattern recognition, pp. 5063\u20135072. https:\/\/doi.org\/10.1109\/CVPR42600.2020.00511","DOI":"10.1109\/CVPR42600.2020.00511"},{"key":"3567_CR22","unstructured":"Memmesheimer R, Theisen N, Paulus D (2020) Signal level deep metric learning for multimodal one-shot action recognition. arXiv: 201213823v1. pp. 1\u20137"},{"key":"3567_CR23","unstructured":"Open-MMLab (2019) mmpose. https:\/\/githubcom\/open-mmlab\/mmpose"},{"key":"3567_CR24","doi-asserted-by":"publisher","unstructured":"Peng W, Hong X, Chen H, Zhao G (2020) Learning graph convolutional network for skeleton-based human action recognition by neural searching. In: AAAI conference on artificial intelligence, New York, USA, pp. 2669\u20132676. https:\/\/doi.org\/10.1609\/aaai.v34i03.5652","DOI":"10.1609\/aaai.v34i03.5652"},{"key":"3567_CR25","doi-asserted-by":"publisher","unstructured":"Qiu Z, Yao T, Mei T (2017) Learning spatio-temporal representation with pseudo-3D residual networks. IEEE international conference on computer vision. pp. 5534\u20135542. https:\/\/doi.org\/10.1109\/ICCV.2017.590","DOI":"10.1109\/ICCV.2017.590"},{"key":"3567_CR26","doi-asserted-by":"crossref","unstructured":"Shahroudy A, Liu J, Ng TT, Wang G (2016) NTU RGB+D: a large scale dataset for 3D human activity analysis. IEEE conference on computer vision and pattern recognition. Las Vegas, USA, pp. 1010\u20131019","DOI":"10.1109\/CVPR.2016.115"},{"key":"3567_CR27","doi-asserted-by":"crossref","unstructured":"Shi L, Zhangng Y, Cheng J, Lu H (2019) Skeleton-based action recognition with directed graph neural networks. IEEE conference on computer vision and pattern recognition. Long Beach, USA, pp. 7912\u20137921","DOI":"10.1109\/CVPR.2019.00810"},{"key":"3567_CR28","doi-asserted-by":"crossref","unstructured":"Si C, Chen W, Wang W, Wang L, Tan T (2019) An attention enhanced graph convolutional LSTM network for skeleton-based action recognition. IEEE\/CVF conference on computer vision and pattern recognition. Los Angeles CA, United States, pp. 1227\u20131236","DOI":"10.1109\/CVPR.2019.00132"},{"key":"3567_CR29","doi-asserted-by":"crossref","unstructured":"Simon T, Joo H, Matthews I, Sheikh Y (2017) Hand keypoint detection in single images using multiview bootstrapping. In: IEEE conference on computer vision and pattern recognition, pp. 1145\u20131153","DOI":"10.1109\/CVPR.2017.494"},{"key":"3567_CR30","unstructured":"Simonyan K (2014) Two-stream convolutional networks for action recognition in videos. 27th International conference on neural information processing systems, pp. 1\u201311, https:\/\/arxiv.org\/pdf\/1406.2199.pdf, arXiv: 1406.2199v2"},{"key":"3567_CR31","doi-asserted-by":"publisher","first-page":"267","DOI":"10.1007\/978-3-319-66808-6_18","volume-title":"Machine learning and knowledge extraction","author":"D Singh","year":"2017","unstructured":"Singh D, Merdivan E, Psychoula I, Kropf J, Hanke S, Geist M, Holzinger A (2017) Human activity recognition using recurrent neural networks. In: Holzinger A, Kieseberg P, Tjoa AM, Weippl E (eds) Machine learning and knowledge extraction. Springer, Cham, pp 267\u2013274"},{"key":"3567_CR32","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-14799-0_53","volume-title":"Violent crowd flow detection using deep learning","author":"SA Sumon","year":"2019","unstructured":"Sumon SA, Shahria MT, Goni MR, Hasan N, Almarufuzzaman AM, Rahman RM (2019) Violent crowd flow detection using deep learning. Springer, Berlin"},{"key":"3567_CR33","doi-asserted-by":"crossref","unstructured":"Takeishi N, Kawahara Y, Yairi T (2017) Learning Koopman invariant subspaces for dynamic mode decomposition. arXiv: 171004340, pp. 1\u201318","DOI":"10.24963\/ijcai.2017\/392"},{"key":"3567_CR34","doi-asserted-by":"publisher","unstructured":"Tran D, Bourdev L, Fergus R, Torresani L, Paluri M (2015) Learning spatiotemporal features with 3D convolutional networks. IEEE international conference on computer vision, pp. 4489\u20134497. https:\/\/doi.org\/10.1109\/ICCV.2015.510","DOI":"10.1109\/ICCV.2015.510"},{"key":"3567_CR35","unstructured":"Tran D, Ray J, Shou Z, Chang SF, Paluri M (2017) Convnet architecture search for spatiotemporal feature learning. arXiv: 170805038, pp. 1\u201310"},{"key":"3567_CR36","doi-asserted-by":"publisher","unstructured":"Wang H, Schmid C (2013) Action recognition with improved trajectories. IEEE international conference on computer vision, pp. 3551\u20133558, https:\/\/doi.org\/10.1109\/ICCV.2013.441","DOI":"10.1109\/ICCV.2013.441"},{"key":"3567_CR37","doi-asserted-by":"crossref","unstructured":"Wang H, Wang L (2017) Modeling temporal dynamics and spatial configurations of actions using two-stream recurrent neural networks. IEEE conference on computer vision and pattern recognition, pp. 499\u2013508","DOI":"10.1109\/CVPR.2017.387"},{"issue":"1","key":"3567_CR38","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1007\/s11263-012-0594-8","volume":"103","author":"H Wang","year":"2013","unstructured":"Wang H, Kl\u00e4ser A, Schmid C, Liu CL (2013) Dense trajectories and motion boundary descriptors for action recognition. Int J Comput Vis 103(1):60\u201379. https:\/\/doi.org\/10.1007\/s11263-012-0594-8","journal-title":"Int J Comput Vis"},{"key":"3567_CR39","doi-asserted-by":"crossref","unstructured":"Wang L, Xiong Y, Wang Z, Qiao Y, Lin D, Tang X, Gool LV (2016) Temporal segment networks: towards good practices for deep action recognition. In: European conference on computer vision, pp. 20\u201336","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"3567_CR40","doi-asserted-by":"crossref","unstructured":"Wei SE, Ramakrishna V, Kanade T, Sheikh Y (2016) Convolutional pose machines. In: IEEE conference on computer vision and pattern recognition, pp. 4724\u20134732","DOI":"10.1109\/CVPR.2016.511"},{"key":"3567_CR41","doi-asserted-by":"crossref","unstructured":"Yan S, Xiong Y, Lin D (2018) Spatial temporal graph convolutional networks for skeleton-based action recognition. In: AAAI conference on artificial intelligence, New Orleans, USA, pp. 1\u201310, arXiv: 1801.07455v2","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"3567_CR42","doi-asserted-by":"crossref","unstructured":"Zhang S, Liu X, Xiao J (2017) On geometric features for skeleton-based action recognition using multilayer LSTM networks. IEEE winter conference on applications of computer vision, pp. 148\u2013157","DOI":"10.1109\/WACV.2017.24"},{"key":"3567_CR43","doi-asserted-by":"publisher","unstructured":"Zhao R, Wang K, Su H, Ji Q (2019) Bayesian graph convolution LSTM for skeleton based action recognition. In: IEEE international conference on computer vision, Los Angeles CA, United States, pp. 6881\u20136891, https:\/\/doi.org\/10.1109\/ICCV.2019.00698","DOI":"10.1109\/ICCV.2019.00698"},{"key":"3567_CR44","doi-asserted-by":"crossref","unstructured":"Zhou B, Andonian A, Oliva A, Torralba A (2018) Temporal relational reasoning in videos. In: European conference on computer vision, pp. 803\u2013818","DOI":"10.1007\/978-3-030-01246-5_49"},{"key":"3567_CR45","unstructured":"Zhu Y, Li X, Liu C, Zolfaghari M, Xiong Y, Wu C, Zhang Z, Tighe J, Manmatha R, Li M (2020) A comprehensive study of deep video action recognition. arXiv: 201206567v1, pp. 1\u201330"}],"container-title":["Journal of Ambient Intelligence and Humanized Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12652-021-03567-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s12652-021-03567-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s12652-021-03567-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,23]],"date-time":"2023-05-23T17:58:01Z","timestamp":1684864681000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s12652-021-03567-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,27]]},"references-count":45,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,6]]}},"alternative-id":["3567"],"URL":"https:\/\/doi.org\/10.1007\/s12652-021-03567-1","relation":{},"ISSN":["1868-5137","1868-5145"],"issn-type":[{"type":"print","value":"1868-5137"},{"type":"electronic","value":"1868-5145"}],"subject":[],"published":{"date-parts":[[2021,10,27]]},"assertion":[{"value":"25 January 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 October 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 October 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The datasets generated during and\/or analysed during the current study are available from the corresponding author on reasonable request.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Data availability"}}]}}