{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,19]],"date-time":"2026-03-19T18:49:37Z","timestamp":1773946177272,"version":"3.50.1"},"reference-count":29,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,5,13]],"date-time":"2023-05-13T00:00:00Z","timestamp":1683936000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,5,13]],"date-time":"2023-05-13T00:00:00Z","timestamp":1683936000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100019065","name":"Tianjin Science and Technology Program","doi-asserted-by":"publisher","award":["No. 20JCYBJC00300."],"award-info":[{"award-number":["No. 20JCYBJC00300."]}],"id":[{"id":"10.13039\/501100019065","id-type":"DOI","asserted-by":"publisher"}]},{"name":"National Education Science Planning","award":["No. BHA220139"],"award-info":[{"award-number":["No. BHA220139"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["EURASIP J. Adv. Signal Process."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Studying the real-time face expression state of teachers in class was important to build an objective classroom teaching evaluation system based on AI. However, the face-to-face communication in classroom conditions was a real-time process that operated on a millisecond time scale. Therefore, in order to quickly and accurately predict teachers\u2019 facial expressions in real time, this paper proposed an improved YOLOv5 network, which introduced the attention mechanisms into the Backbone model of YOLOv5. In experiments, we investigated the effects of different attention mechanisms on YOLOv5 by adding different attention mechanisms after each CBS module in the CSP1_X structure of the Backbone part, respectively. At the same time, the attention mechanisms were incorporated at different locations of the Focus, CBS, and SPP modules of YOLOv5, respectively, to study the effects of the attention mechanism on different modules. The results showed that the network in which the coordinate attentions were incorporated after each CBS module in the CSP1_X structure obtained the detection time of 25\u00a0ms and the accuracy of 77.1% which increased by 3.5% compared with YOLOv5. It outperformed other networks, including Faster-RCNN, R-FCN, ResNext-101, DETR, Swin-Transformer, YOLOv3, and YOLOX. Finally, the real-time teachers\u2019 facial expression recognition system was designed to detect and analyze the teachers\u2019 facial expression distribution with time through camera and the teaching video.<\/jats:p>","DOI":"10.1186\/s13634-023-01019-w","type":"journal-article","created":{"date-parts":[[2023,5,13]],"date-time":"2023-05-13T10:02:16Z","timestamp":1683972136000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Research on real-time teachers\u2019 facial expression recognition based on YOLOv5 and attention mechanisms"],"prefix":"10.1186","volume":"2023","author":[{"given":"Hongmei","family":"Zhong","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8948-451X","authenticated-orcid":false,"given":"Tingting","family":"Han","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Xia","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yan","family":"Tian","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Libao","family":"Wu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,5,13]]},"reference":[{"key":"1019_CR1","doi-asserted-by":"publisher","unstructured":"P.P. Filntisis, A. Katsamanis, P. Maragos, Photorealistic adaptation and interpolation of facial expressions using HMMS and AAMS for audio-visual speech synthesis, in IEEE International Conference on Image Processing (ICIP), Beijing, China, pp. 2941\u20132945. https:\/\/doi.org\/10.1109\/ICIP.2017.8296821 (2017).","DOI":"10.1109\/ICIP.2017.8296821"},{"key":"1019_CR2","doi-asserted-by":"publisher","unstructured":"A. Halder, A. Chakraborty, A. Konar, A. K. Nagar, Computing with words model for emotion recognition by facial expression analysis using interval type-2 fuzzy sets, in IEEE International Conference on Fuzzy Systems (FUZZ-IEEE), Hyderabad, India, pp. 1\u20138. https:\/\/doi.org\/10.1109\/FUZZ-IEEE.2013.6622543 (2013).","DOI":"10.1109\/FUZZ-IEEE.2013.6622543"},{"key":"1019_CR3","doi-asserted-by":"publisher","unstructured":"N. Sebe, M.S. Lew, I. Cohen, A. Garg, T.S. Huang, Emotion recognition using a Cauchy Naive Bayes classifier, in International Conference on Pattern Recognition, Quebec City, QC, Canada, pp. 17\u201320. https:\/\/doi.org\/10.1109\/ICPR.2002.1044578 (2002).","DOI":"10.1109\/ICPR.2002.1044578"},{"key":"1019_CR4","doi-asserted-by":"publisher","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","volume":"39","author":"S Ren","year":"2017","unstructured":"S. Ren, K. He, R. Girshick, J. Sun, Faster R-CNN: towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 39, 1137\u20131149 (2017). https:\/\/doi.org\/10.1109\/TPAMI.2016.2577031","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"1019_CR5","unstructured":"J. Dai, Y. Li, K. He, et al., R-FCN: Object Detection via Region-based Fully Convolutional Networks (Curran Associates Inc, 2016), p. 379\u2013387."},{"key":"1019_CR6","doi-asserted-by":"crossref","unstructured":"W. Liu, et al. SSD: Single Shot MultiBox Detector. European Conference on Computer Vision. arXiv preprint, arXiv:1512.02325 (2016).","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"1019_CR7","doi-asserted-by":"publisher","unstructured":"J. Redmon, S. Divvala, R. Girshick, A. Farhadi, You only look once: unified, real-time object detection, in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, pp. 779\u2013788. https:\/\/doi.org\/10.1109\/CVPR.2016.91 (2016).","DOI":"10.1109\/CVPR.2016.91"},{"key":"1019_CR8","unstructured":"Simonyan, K., & Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition.\u00a0arXiv preprint, arXiv:1409.1556 (2014)."},{"key":"1019_CR9","doi-asserted-by":"publisher","unstructured":"C. Szegedy et al., Going deeper with convolutions, in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, pp. 1\u20139. https:\/\/doi.org\/10.1109\/CVPR.2015.7298594 (2015).","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"1019_CR10","doi-asserted-by":"publisher","unstructured":"H. Jun, L. Shuai, S. Jinming, L. Yue, W. Jingwei, J. Peng, Facial expression recognition based on VGGNet convolutional neural network, in Chinese Automation Congress (CAC), Xi'an, China, pp. 4146\u20134151. https:\/\/doi.org\/10.1109\/CAC.2018.8623238 (2018).","DOI":"10.1109\/CAC.2018.8623238"},{"key":"1019_CR11","unstructured":"A. Khanzada, C. Bai, F. T. Celepcikay, Facial Expression Recognition with Deep Learning, arXiv preprint, arXiv:2004.11823 (2020)."},{"issue":"1","key":"1019_CR12","first-page":"210","volume":"47","author":"LX bing","year":"2021","unstructured":"L. X. bing, C. Lian, Face detection in natural scene based on improved Faster-RCNN. Comput. Eng. 47(1), 210\u2013216 (2021).","journal-title":"Comput. Eng"},{"issue":"2022","key":"1019_CR13","doi-asserted-by":"publisher","first-page":"106694","DOI":"10.1016\/j.compag.2022.106694","volume":"193","author":"AM Roy","year":"2022","unstructured":"A.M. Roy, J. Bhaduri, Real-time growth stage detection model for high degree of occultation using DenseNet-fused YOLOv4. Comput. Electron. Agric. 193(2022), 106694 (2022). https:\/\/doi.org\/10.1016\/j.compag.2022.106694","journal-title":"Comput. Electron. Agric."},{"key":"1019_CR14","doi-asserted-by":"publisher","first-page":"1447","DOI":"10.1038\/s41598-021-81216-5","volume":"11","author":"MO Lawal","year":"2021","unstructured":"M.O. Lawal, Tomato detection based on modified YOLOv3 framework. Sci. Rep. 11, 1447 (2021). https:\/\/doi.org\/10.1038\/s41598-021-81216-5","journal-title":"Sci. Rep."},{"key":"1019_CR15","doi-asserted-by":"publisher","first-page":"3895","DOI":"10.1007\/s00521-021-06651-x","volume":"34","author":"AM Roy","year":"2022","unstructured":"A.M. Roy, R. Bose, J. Bhaduri, A fast accurate fine-grain object detection model based on YOLOv4 deep neural network. Neural Comput. Appl. 34, 3895\u20133921 (2022). https:\/\/doi.org\/10.1007\/s00521-021-06651-x","journal-title":"Neural Comput. Appl."},{"key":"1019_CR16","doi-asserted-by":"publisher","unstructured":"H. Aung, A. V. Bobkov and N. L. Tun, Face detection in real time live video using yolo algorithm based on Vgg16 convolutional neural network, in International Conference on Industrial Engineering, Applications and Manufacturing (ICIEAM), Sochi, Russia, pp. 697\u2013702. https:\/\/doi.org\/10.1109\/ICIEAM51226.2021.9446291 (2021).","DOI":"10.1109\/ICIEAM51226.2021.9446291"},{"key":"1019_CR17","doi-asserted-by":"publisher","unstructured":"J. Redmon, A. Farhadi, YOLO9000: better, faster, stronger, in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, pp. 6517\u20136525. https:\/\/doi.org\/10.1109\/CVPR.2017.690 (2017).","DOI":"10.1109\/CVPR.2017.690"},{"key":"1019_CR18","unstructured":"J. Redmon, A. Farhadi. YOLOv3: An Incremental Improvement. arXiv preprint, arXiv:1804.02767 (2018)."},{"key":"1019_CR19","unstructured":"B. Alexey, W. Chien-Yao, Y.M.L. Hong, YOLOv4: optimal speed and accuracy of object detection.\u00a0arXiv preprint, arXiv:2004.10934v1 (2020)."},{"key":"1019_CR20","doi-asserted-by":"crossref","unstructured":"Y. Fan, X. Lu, D. Li, et al., Video-based emotion recognition using CNN-RNN and C3D hybrid networks, in Proceedings of the 18th ACM International Conference on Multimodal Interaction, pp. 445\u2013450 (2016).","DOI":"10.1145\/2993148.2997632"},{"key":"1019_CR21","doi-asserted-by":"publisher","unstructured":"Liu, Research and Implementation of Face Expression Recognition Algorithm based on Video Image. CQUPT, https:\/\/doi.org\/10.27675\/d.cnki.gcydx.2021.001271 (2021).","DOI":"10.27675\/d.cnki.gcydx.2021.001271"},{"key":"1019_CR22","doi-asserted-by":"publisher","unstructured":"Zhang, Video Emotion Recognition Based on Dual-stream Network. JLU, https:\/\/doi.org\/10.27162\/d.cnki.gjlin.2022.002308 (2022).","DOI":"10.27162\/d.cnki.gjlin.2022.002308"},{"issue":"2","key":"1019_CR23","first-page":"45","volume":"42","author":"QX Fei","year":"2020","unstructured":"Q.X. Fei, S. Kai, Z. Yue, Y. Yong, Z. Gang, J. Cheng, L.C. Ming, L.X. Dong, Z.J. Feng, Detection algorithm for key points on face based on attention model. Opt. Instrum. 42(2), 45\u201349 (2020)","journal-title":"Opt. Instrum."},{"issue":"04","key":"1019_CR24","first-page":"159","volume":"38","author":"J Kang","year":"2020","unstructured":"J. Kang, S. Li, Convolutional neural network face expression recognition based on attention mechanism. J. Shaanxi Univ. Sci. Technol. 38(04), 159\u2013165+171 (2020)","journal-title":"J. Shaanxi Univ. Sci. Technol."},{"key":"1019_CR25","doi-asserted-by":"publisher","first-page":"2011","DOI":"10.1109\/TPAMI.2019.2913372","volume":"42","author":"J Hu","year":"2020","unstructured":"J. Hu, L. Shen, S. Albanie, G. Sun, E. Wu, Squeeze-and-excitation networks. IEEE Trans. Pattern Anal. Mach. Intell. 42, 2011\u20132023 (2020). https:\/\/doi.org\/10.1109\/TPAMI.2019.2913372","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"1019_CR26","doi-asserted-by":"publisher","unstructured":"Q. Wang, B. Wu, P. Zhu, P. Li, W. Zuo, Q. Hu, ECA-Net: efficient channel attention for deep convolutional neural networks, in IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, pp. 11531\u201311539. https:\/\/doi.org\/10.1109\/CVPR42600.2020.01155 (2020).","DOI":"10.1109\/CVPR42600.2020.01155"},{"key":"1019_CR27","doi-asserted-by":"crossref","unstructured":"S. Woo, J. Park, J.Y. Lee, I.S. Kweon, CBAM: Convolutional Block Attention Module. arXiv preprint, arXiv:1807.06521v2 (2018).","DOI":"10.1007\/978-3-030-01234-2_1"},{"issue":"1","key":"1019_CR28","doi-asserted-by":"publisher","first-page":"273","DOI":"10.1049\/ipr2.12364","volume":"16","author":"C Xie","year":"2022","unstructured":"C. Xie, H. Zhu, Y. Fei, Deep coordinate attention network for single image super-resolution. IET Image Proc. 16(1), 273\u2013284 (2022)","journal-title":"IET Image Proc."},{"key":"1019_CR29","doi-asserted-by":"publisher","unstructured":"S. Li, W. Deng and J. Du, Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild, in \u2013IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, pp. 2584\u20132593. https:\/\/doi.org\/10.1109\/CVPR.2017.277 (2017).","DOI":"10.1109\/CVPR.2017.277"}],"container-title":["EURASIP Journal on Advances in Signal Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13634-023-01019-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13634-023-01019-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13634-023-01019-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,13]],"date-time":"2023-05-13T10:03:14Z","timestamp":1683972194000},"score":1,"resource":{"primary":{"URL":"https:\/\/asp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13634-023-01019-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,13]]},"references-count":29,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["1019"],"URL":"https:\/\/doi.org\/10.1186\/s13634-023-01019-w","relation":{},"ISSN":["1687-6180"],"issn-type":[{"value":"1687-6180","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,13]]},"assertion":[{"value":"3 September 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 May 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 May 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The article has been ethically approved and approved.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"All presentations of the case report have been agreed for publication.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"55"}}