{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T06:12:29Z","timestamp":1778825549015,"version":"3.51.4"},"reference-count":24,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2020,3,12]],"date-time":"2020-03-12T00:00:00Z","timestamp":1583971200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,3,12]],"date-time":"2020-03-12T00:00:00Z","timestamp":1583971200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100014219","name":"National Science Fund for Distinguished Young Scholars","doi-asserted-by":"crossref","award":["No.61425002"],"award-info":[{"award-number":["No.61425002"]}],"id":[{"id":"10.13039\/501100014219","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100018621","name":"Program for ChangJiang Scholars and Innovative Research Team in University","doi-asserted-by":"crossref","award":["No.IRT_15R07"],"award-info":[{"award-number":["No.IRT_15R07"]}],"id":[{"id":"10.13039\/501100018621","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100013099","name":"Scientific Research Fund of Liaoning Provincial Education Department","doi-asserted-by":"publisher","award":["No.L2019606"],"award-info":[{"award-number":["No.L2019606"]}],"id":[{"id":"10.13039\/501100013099","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"e National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["Nos. 91748104, 61632006, 61877008"],"award-info":[{"award-number":["Nos. 91748104, 61632006, 61877008"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Program for the Liaoning Distinguished Professor, Program for Dalian High-level Talent Innovation Support","award":["No.2017RD11"],"award-info":[{"award-number":["No.2017RD11"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Comput. Ind. Biomed. Art"],"published-print":{"date-parts":[[2020,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>With the rapid development of deep learning technology, behavior recognition based on video streams has made great progress in recent years. However, there are also some problems that must be solved: (1) In order to improve behavior recognition performance, the models have tended to become deeper, wider, and more complex. However, some new problems have been introduced also, such as that their real-time performance decreases; (2) Some actions in existing datasets are so similar that they are difficult to distinguish. To solve these problems, the ResNet34-3DRes18 model, which is a lightweight and efficient two-dimensional (2D) and three-dimensional (3D) fused model, is constructed in this study. The model used 2D convolutional neural network (2DCNN) to obtain the feature maps of input images and 3D convolutional neural network (3DCNN) to process the temporal relationships between frames, which made the model not only make use of 3DCNN\u2019s advantages on video temporal modeling but reduced model complexity. Compared with state-of-the-art models, this method has shown excellent performance at a faster speed. Furthermore, to distinguish between similar motions in the datasets, an attention gate mechanism is added, and a Res34-SE-IM-Net attention recognition model is constructed. The Res34-SE-IM-Net achieved 71.85%, 92.196%, and 36.5% top-1 accuracy (The predicting label obtained from model is the largest one in the output probability vector. If the label is the same as the target label of the motion, the classification is correct.) respectively on the test sets of the HMDB51, UCF101, and Something-Something v1 datasets.<\/jats:p>","DOI":"10.1186\/s42492-020-00045-x","type":"journal-article","created":{"date-parts":[[2020,3,12]],"date-time":"2020-03-12T00:07:48Z","timestamp":1583971668000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["Fused behavior recognition model based on attention mechanism"],"prefix":"10.1186","volume":"3","author":[{"given":"Lei","family":"Chen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5200-640X","authenticated-orcid":false,"given":"Rui","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dongsheng","family":"Zhou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xin","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qiang","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2020,3,12]]},"reference":[{"key":"45_CR1","doi-asserted-by":"publisher","unstructured":"Kuehne H, Jhuang H, Garrote E, Poggio T, Serre T (2011) HMDB: a large video database for human motion recognition. Paper presented at 2011 IEEE international conference on computer vision, IEEE, Barcelona, pp. 2556\u20132563 https:\/\/doi.org\/10.1109\/ICCV.2011.6126543","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"45_CR2","unstructured":"Somro K, Zamir AR, Shah M (2012) UCF101: a dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012"},{"key":"45_CR3","unstructured":"Kay W, Carreira J, Simonyan K, Zhang B, Hillier C, Vijayanarasimhan S, et al (2017) The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017"},{"key":"45_CR4","doi-asserted-by":"publisher","unstructured":"Goyal R, Kahou SE, Michalski V, Materzynska J, Westphal S, Kim H et al (2017) The \u201csomething something\u201d video database for learning and evaluating visual common sense. Paper presented at 2017 IEEE international conference on computer vision, IEEE, Venice, pp. 5843\u20135851 https:\/\/doi.org\/10.1109\/ICCV.2017.622","DOI":"10.1109\/ICCV.2017.622"},{"key":"45_CR5","first-page":"568","volume":"2014","author":"K Simonyan","year":"2014","unstructured":"Simonyan K, Zisserman A (2014) Two-stream convolutional networks for action recognition in videos. Adv Neural Inf Proces Syst 2014:568\u2013576","journal-title":"Adv Neural Inf Proces Syst"},{"key":"45_CR6","doi-asserted-by":"publisher","unstructured":"Feichtenhofer C, Pinz A, Zisserman A (2016) Convolutional two-stream network fusion for video action recognition. Paper presented at the 29th IEEE conference on computer vision and pattern recognition, IEEE, Las Vegas, pp. 1933\u20131941 https:\/\/doi.org\/10.1109\/CVPR.2016.213","DOI":"10.1109\/CVPR.2016.213"},{"key":"45_CR7","doi-asserted-by":"publisher","first-page":"20","DOI":"10.1007\/978-3-319-46484-8_2","volume-title":"Computer vision \u2013 ECCV 2016. Paper presented at the 14th European conference on computer vision ECCV, lecture notes in computer science","author":"LM Wang","year":"2016","unstructured":"Wang LM, Xiong YJ, Wang Z, Qiao Y, Lin DH, Tang XO et al (2016) Temporal segment networks: towards good practices for deep action recognition. In: Leibe B, Matas J, Sebe N, Welling M (eds) Computer vision \u2013 ECCV 2016. Paper presented at the 14th European conference on computer vision ECCV, lecture notes in computer science, vol 9912. Springer, Cham, pp 20\u201336. https:\/\/doi.org\/10.1007\/978-3-319-46484-8_2"},{"key":"45_CR8","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510","volume-title":"Learning spatiotemporal features with 3D convolutional networks","author":"D Tran","year":"2015","unstructured":"Tran D, Bourdev L, Fergus R, Torresani L, Paluri M (2015) Learning spatiotemporal features with 3D convolutional networks. Paper presented at 2015 IEEE international conference on computer vision, IEEE, Santiago Chile, 7-13 December 2015. https:\/\/doi.org\/10.1109\/ICCV.2015.510"},{"key":"45_CR9","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502","volume-title":"Quo vadis, action recognition? A new model and the kinetics dataset","author":"J Carreira","year":"2017","unstructured":"Carreira J, Zisserman A (2017) Quo vadis, action recognition? A new model and the kinetics dataset. Paper presented at 2017 IEEE conference on computer vision and pattern recognition, IEEE, Honolulu Hawaii, 21-26 July 2017. https:\/\/doi.org\/10.1109\/CVPR.2017.502"},{"key":"45_CR10","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.590","volume-title":"Learning Spatio-temporal representation with pseudo-3D residual networks","author":"ZF Qiao","year":"2017","unstructured":"Qiao ZF, Yao T, Mei T (2017) Learning Spatio-temporal representation with pseudo-3D residual networks. Paper presented at 2017 IEEE international conference on computer vision, IEEE, Venice, 22-29 October 2017. https:\/\/doi.org\/10.1109\/ICCV.2017.590"},{"key":"45_CR11","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00675","volume-title":"A closer look at Spatio-temporal convolutions for action recognition","author":"D Tran","year":"2018","unstructured":"Tran D, Wang H, Torresani L, Ray J, LeCun Y, Paluri M (2018) A closer look at Spatio-temporal convolutions for action recognition. Paper presented at the 31th IEEE conference on computer vision and pattern recognition, IEEE, salt Lake, 18-23 June 2018. https:\/\/doi.org\/10.1109\/CVPR.2018.00675"},{"key":"45_CR12","doi-asserted-by":"publisher","first-page":"8","DOI":"10.1007\/978-3-030-01216-8_43","volume-title":"Proceedings of 15th European conference on computer vision","author":"M Zolfaghari","year":"2018","unstructured":"Zolfaghari M, Singh K (2018) Brox T (2018) ECO: efficient convolutional network for online video understanding. In: Ferrari V, Hebert M, Sminchisescu C, Weiss Y (eds) Proceedings of 15th European conference on computer vision. Springer, Cham, pp 8\u201314. https:\/\/doi.org\/10.1007\/978-3-030-01216-8_43"},{"key":"45_CR13","volume-title":"TSM: temporal shift module for efficient video understanding","author":"J Lin","year":"2019","unstructured":"Lin J, Gan C, Hang S (2019) TSM: temporal shift module for efficient video understanding. Paper presented at 2019 IEEE international conference on computer vision, IEEE, Seoul Korea, 27 October-3 November 2019"},{"key":"45_CR14","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90","volume-title":"Deep residual learning for image recognition","author":"KM He","year":"2016","unstructured":"He KM, Zhang XY, Ren SQ, Sun J (2016) Deep residual learning for image recognition. Paper presented at the 2016 IEEE conference on computer vision and pattern recognition, IEEE, Las Vegas, 27-30 June 2016. https:\/\/doi.org\/10.1109\/CVPR.2016.90"},{"key":"45_CR15","unstructured":"Tran D, Ray J, Shou Z, Chang SF, Paluri M (2017) Convnet architecture search for spatiotemporal feature learning. arXiv preprint arXiv:1708.05038, 2017"},{"key":"45_CR16","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00745","volume-title":"Squeeze-and-excitation networks","author":"J Hu","year":"2018","unstructured":"Hu J, Shen L, Sun G (2018) Squeeze-and-excitation networks. Paper presented at the 2018 IEEE conference on computer vision and pattern recognition, IEEE, salt Lake, 18-23 June 2018. https:\/\/doi.org\/10.1109\/CVPR.2018.00745"},{"key":"45_CR17","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.441","volume-title":"Action recognition with improved trajectories","author":"H Wang","year":"2013","unstructured":"Wang H, Schmid C (2013) Action recognition with improved trajectories. Paper presented at paper presented at 2013 IEEE international conference on computer vision, IEEE, Sydney, 1-8 December 2013. https:\/\/doi.org\/10.1109\/ICCV.2013.441"},{"key":"45_CR18","volume-title":"Beyond Gaussian pyramid: multi-skip feature stacking for action recognition","author":"ZZ Lan","year":"2015","unstructured":"Lan ZZ, Lin M, Li XC, Al G, Raj B 2015 Beyond Gaussian pyramid: multi-skip feature stacking for action recognition. Paper presented at the 2015 IEEE conference on computer vision and pattern recognition, IEEE, Boston, 7-12 June 2015"},{"key":"45_CR19","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00685","volume-title":"can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet?","author":"K Hara","year":"2018","unstructured":"Hara K, Kataoka H, Satoh Y (2018) can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet? Paper presented at the 2018 IEEE conference on computer vision and pattern recognition, IEEE, salt Lake, 18\u201323 2018. https:\/\/doi.org\/10.1109\/CVPR.2018.00685"},{"key":"45_CR20","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2018.8545710","volume-title":"End-to-end video-level representation learning for action recognition","author":"JG Zhu","year":"2018","unstructured":"Zhu JG, Zou W, Zhu Z (2018) End-to-end video-level representation learning for action recognition. Paper presented at the 24th international conference on pattern recognition, IEEE, Beijing China, 20-24 august 2018. https:\/\/doi.org\/10.1109\/ICPR.2018.8545710"},{"key":"45_CR21","doi-asserted-by":"publisher","first-page":"111043","DOI":"10.1109\/ACCESS.2019.2933303","volume":"7","author":"Chunlei Wu","year":"2019","unstructured":"Wu CL, Cao HW, Zhang WS, Wang LQ, Wei YW, Peng ZX (2019) Refined spatial network for human action recognition. IEEE Access (7):111043\u2013111052. https:\/\/doi.org\/10.1109\/ACCESS.2019.2933303","journal-title":"IEEE Access"},{"key":"45_CR22","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33019167","volume-title":"Memory-augmented temporal dynamic learning for action recognition","author":"Y Yuan","year":"2019","unstructured":"Yuan Y, Wang D, Wang Q (2019) Memory-augmented temporal dynamic learning for action recognition. Paper presented at the 33th AAAI conference on artificial intelligence. 33: 9167-9175. https:\/\/doi.org\/10.1609\/aaai.v33i01.33019167"},{"key":"45_CR23","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_49","volume-title":"Temporal relational reasoning in videos","author":"BL Zhou","year":"2018","unstructured":"Zhou BL, Andonian A, Oliva A, Torralba A (2018) Temporal relational reasoning in videos. Paper presented at the 15th European conference on computer vision, springer, Munich, 8-14 September 2018. https:\/\/doi.org\/10.1007\/978-3-030-01246-5_49"},{"key":"45_CR24","unstructured":"Shi P (2018) Research of speech emotion recognition based on deep neural network. Dissertation, Wuhan University of Technology"}],"container-title":["Visual Computing for Industry, Biomedicine, and Art"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s42492-020-00045-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s42492-020-00045-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s42492-020-00045-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,3,12]],"date-time":"2021-03-12T00:07:52Z","timestamp":1615507672000},"score":1,"resource":{"primary":{"URL":"https:\/\/vciba.springeropen.com\/articles\/10.1186\/s42492-020-00045-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,12]]},"references-count":24,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,12]]}},"alternative-id":["45"],"URL":"https:\/\/doi.org\/10.1186\/s42492-020-00045-x","relation":{},"ISSN":["2524-4442"],"issn-type":[{"value":"2524-4442","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,3,12]]},"assertion":[{"value":"12 December 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 February 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"12 March 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The content of Fig.  comes from the public database. The role in Fig.  is one of the authors for this article. This article does not infringe the right of portrait.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"7"}}