{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,28]],"date-time":"2026-07-28T04:03:35Z","timestamp":1785211415063,"version":"3.55.0"},"reference-count":23,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2020,6,17]],"date-time":"2020-06-17T00:00:00Z","timestamp":1592352000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,6,17]],"date-time":"2020-06-17T00:00:00Z","timestamp":1592352000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Image Video Proc."],"published-print":{"date-parts":[[2020,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper addresses the recognitions of human actions in videos. Human action recognition can be seen as the automatic labeling of a video according to the actions occurring in it. It has become one of the most challenging and attractive problems in the pattern recognition and video classification fields. The problem itself is difficult to solve by traditional video processing methods because of several challenges such as the background noise, sizes of subjects in different videos, and the speed of actions. Derived from the progress of deep learning methods, several directions are developed to recognize a human action from a video, such as the long-short-term memory (LSTM)-based model, two-stream convolutional neural network (CNN) model, and the convolutional 3D model.In this paper, we focus on the two-stream structure. The traditional two-stream CNN network solves the problem that CNNs do not have satisfactory performance on temporal features. By training a temporal stream, which uses the optical flow as the input, a CNN can have the ability to extract temporal features. However, the optical flow only contains limited temporal information because it only records the movements of pixels on the<jats:italic>x<\/jats:italic>-axis and the<jats:italic>y<\/jats:italic>-axis. Therefore, we attempt to design and implement a new two-stream model by using an LSTM-based model in its spatial stream to extract both spatial and temporal features in RGB frames. In addition, we implement a DenseNet in the temporal stream to improve the recognition accuracy. This is in-contrast to traditional approaches which typically utilize the spatial stream for extracting only spatial features. The quantitative evaluation and experiments are conducted on the UCF-101 dataset, which is a well-developed public video dataset. For the temporal stream, we choose the optical flow of UCF-101. Images in the optical flow are provided by the Graz University of Technology. The experimental result shows that the proposed method outperforms the traditional two-stream CNN method with an accuracy of at least 3%. For both spatial and temporal streams, the proposed model also achieves higher recognition accuracies. In addition, compared with the state of the art methods, the new model can still have the best recognition performance.<\/jats:p>","DOI":"10.1186\/s13640-020-00501-x","type":"journal-article","created":{"date-parts":[[2020,6,17]],"date-time":"2020-06-17T12:02:58Z","timestamp":1592395378000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":58,"title":["Improved two-stream model for human action recognition"],"prefix":"10.1186","volume":"2020","author":[{"given":"Yuxuan","family":"Zhao","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ka Lok","family":"Man","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jeremy","family":"Smith","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kamran","family":"Siddique","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sheng-Uei","family":"Guan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,6,17]]},"reference":[{"issue":"2","key":"501_CR1","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1016\/j.cviu.2004.02.005","volume":"96","author":"S. Hongeng","year":"2004","unstructured":"S. Hongeng, R. Nevatia, F. Bremond, Video-based event recognition: activity representation and probabilistic recognition methods. Comput. Vis. Image Underst.96(2), 129\u2013162 (2004).","journal-title":"Comput. Vis. Image Underst."},{"issue":"5","key":"501_CR2","doi-asserted-by":"publisher","first-page":"1005","DOI":"10.3390\/s19051005","volume":"19","author":"H. -B. Zhang","year":"2019","unstructured":"H. -B. Zhang, Y. -X. Zhang, B. Zhong, Q. Lei, L. Yang, J. -X. Du, D. -S. Chen, A comprehensive survey of vision-based human action recognition methods. Sensors. 19(5), 1005 (2019).","journal-title":"Sensors"},{"key":"501_CR3","doi-asserted-by":"publisher","unstructured":"H. Jhuang, T. Serre, L. Wolf, T. Poggio, in 2007 IEEE 11th International Conference on Computer Vision. A biologically inspired system for action recognition (IEEE, 2007), pp. 1\u20138. https:\/\/doi.org\/10.1109\/iccv.2007.4408988.","DOI":"10.1109\/iccv.2007.4408988"},{"key":"501_CR4","doi-asserted-by":"publisher","unstructured":"H. Wang, C. Schmid, in Proceedings of the IEEE International Conference on Computer Vision. Action recognition with improved trajectories, (2013), pp. 3551\u20133558. https:\/\/doi.org\/10.1109\/iccv.2013.441.","DOI":"10.1109\/iccv.2013.441"},{"issue":"1","key":"501_CR5","doi-asserted-by":"publisher","first-page":"221","DOI":"10.1109\/TPAMI.2012.59","volume":"35","author":"S. Ji","year":"2012","unstructured":"S. Ji, W. Xu, M. Yang, K. Yu, 3D convolutional neural networks for human action recognition. IEEE Trans. Pattern Anal. Mach. Intell.35(1), 221\u2013231 (2012).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"501_CR6","doi-asserted-by":"publisher","unstructured":"A. Krizhevsky, I. Sutskever, G. E. Hinton, in Advances in Neural Information Processing Systems. ImageNet classification with deep convolutional neural networks, (2012), pp. 1097\u20131105. https:\/\/doi.org\/10.1145\/3065386.","DOI":"10.1145\/3065386"},{"key":"501_CR7","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1109\/TPAMI.2011.157","volume":"3","author":"Z. Zhang","year":"2012","unstructured":"Z. Zhang, D. Tao, Slow feature analysis for human action recognition. IEEE Trans. Pattern Anal. Mach. Intell.3:, 436\u2013450 (2012). https:\/\/doi.org\/10.1109\/tpami.2011.157.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"501_CR8","doi-asserted-by":"publisher","unstructured":"J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, T. Darrell, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Long-term recurrent convolutional networks for visual recognition and description, (2015), pp. 2625\u20132634. https:\/\/doi.org\/10.21236\/ada623249.","DOI":"10.21236\/ada623249"},{"key":"501_CR9","unstructured":"K. Simonyan, A. Zisserman, in Advances in Neural Information Processing Systems. Two-stream convolutional networks for action recognition in videos, (2014), pp. 568\u2013576."},{"issue":"1-2","key":"501_CR10","doi-asserted-by":"publisher","first-page":"221","DOI":"10.1016\/S0925-2312(03)00375-8","volume":"55","author":"C. Gold","year":"2003","unstructured":"C. Gold, P. Sollich, Model selection for support vector machine classification. Neurocomputing. 55(1-2), 221\u2013249 (2003).","journal-title":"Neurocomputing"},{"key":"501_CR11","unstructured":"K. Simonyan, A. Zisserman, Very deep convolutional networks for large-scale image recognition. arXiv preprint (2014). arXiv:1409.1556."},{"key":"501_CR12","doi-asserted-by":"publisher","unstructured":"J. Deng, W. Dong, R. Socher, L. -J. Li, K. Li, L. Fei-Fei, in 2009 IEEE Conference on Computer Vision and Pattern Recognition. ImageNet: a large-scale hierarchical image database (IEEE, 2009), pp. 248\u2013255. https:\/\/doi.org\/10.1109\/cvpr.2009.5206848.","DOI":"10.1109\/cvpr.2009.5206848"},{"issue":"8","key":"501_CR13","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S. Hochreiter","year":"1997","unstructured":"S. Hochreiter, J. Schmidhuber, Long short-term memory. Neural Comput.9(8), 1735\u20131780 (1997).","journal-title":"Neural Comput."},{"key":"501_CR14","doi-asserted-by":"publisher","unstructured":"G. Huang, Z. Liu, L. Van Der Maaten, K. Q. Weinberger, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Densely connected convolutional networks, (2017), pp. 4700\u20134708. https:\/\/doi.org\/10.1109\/cvpr.2017.243.","DOI":"10.1109\/cvpr.2017.243"},{"key":"501_CR15","unstructured":"K. Soomro, A. R. Zamir, M. Shah, Ucf101: a dataset of 101 human actions classes from videos in the wild. arXiv preprint (2012). arXiv:1212.0402."},{"key":"501_CR16","doi-asserted-by":"publisher","unstructured":"X. Xia, C. Xu, B. Nan, in 2017 2nd International Conference on Image, Vision and Computing (ICIVC). Inception-v3 for flower classification (IEEE, 2017), pp. 783\u2013787. https:\/\/doi.org\/10.1109\/icivc.2017.7984661.","DOI":"10.1109\/icivc.2017.7984661"},{"key":"501_CR17","doi-asserted-by":"crossref","unstructured":"C. Szegedy, S. Ioffe, V. Vanhoucke, A. A. Alemi, in Thirty-First AAAI Conference on Artificial Intelligence. Inception-v4, Inception-ResNet and the impact of residual connections on learning, (2017).","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"501_CR18","doi-asserted-by":"publisher","unstructured":"D. Tran, L. Bourdev, R. Fergus, L. Torresani, M. Paluri, in Proceedings of the IEEE International Conference on Computer Vision. Learning spatiotemporal features with 3D convolutional networks, (2015), pp. 4489\u20134497. https:\/\/doi.org\/10.1109\/iccv.2015.510.","DOI":"10.1109\/iccv.2015.510"},{"key":"501_CR19","doi-asserted-by":"publisher","unstructured":"J. Carreira, A. Zisserman, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Quo vadis, action recognition? A new model and the kinetics dataset, (2017), pp. 6299\u20136308. https:\/\/doi.org\/10.1109\/cvpr.2017.502.","DOI":"10.1109\/cvpr.2017.502"},{"key":"501_CR20","doi-asserted-by":"publisher","unstructured":"L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, L. Van Gool, in European Conference on Computer Vision. Temporal segment networks: towards good practices for deep action recognition (Springer, 2016), pp. 20\u201336. https:\/\/doi.org\/10.1007\/978-3-319-46484-8_2.","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"501_CR21","doi-asserted-by":"crossref","unstructured":"Z. Hu, E. -J. Lee, in 2019 IEEE International Conference on Computation, Communication and Engineering (ICCCE). Human motion recognition based on improved 3-dimensional convolutional neural network (IEEE, 2019), pp. 154\u2013156.","DOI":"10.1109\/ICCCE48422.2019.9010816"},{"key":"501_CR22","doi-asserted-by":"publisher","first-page":"16639","DOI":"10.1109\/ACCESS.2018.2814075","volume":"6","author":"A. Dilawari","year":"2018","unstructured":"A. Dilawari, M. U. G. Khan, A. Farooq, Z. -U. Rehman, S. Rho, I. Mehmood, Natural language description of video streams using task-specific feature encoding. IEEE Access. 6:, 16639\u201316645 (2018).","journal-title":"IEEE Access"},{"key":"501_CR23","doi-asserted-by":"publisher","first-page":"16","DOI":"10.1016\/j.compeleceng.2016.06.013","volume":"54","author":"S. Kang","year":"2016","unstructured":"S. Kang, W. Ji, S. Rho, V. A. Padigala, Y. Chen, Cooperative mobile video transmission for traffic surveillance in smart cities. Comput. Electr. Eng.54:, 16\u201325 (2016).","journal-title":"Comput. Electr. Eng."}],"container-title":["EURASIP Journal on Image and Video Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13640-020-00501-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13640-020-00501-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13640-020-00501-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,10,29]],"date-time":"2022-10-29T01:12:20Z","timestamp":1667005940000},"score":1,"resource":{"primary":{"URL":"https:\/\/jivp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13640-020-00501-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,6,17]]},"references-count":23,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,12]]}},"alternative-id":["501"],"URL":"https:\/\/doi.org\/10.1186\/s13640-020-00501-x","relation":{},"ISSN":["1687-5281"],"issn-type":[{"value":"1687-5281","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,6,17]]},"assertion":[{"value":"8 November 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 March 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 June 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare that they have no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"24"}}