{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T20:25:15Z","timestamp":1784924715375,"version":"3.55.0"},"reference-count":41,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T00:00:00Z","timestamp":1761523200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Henan Provincial Research and Practice Project on Higher Education Teaching Reform","award":["2024JGLX0469"],"award-info":[{"award-number":["2024JGLX0469"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Accurately recognizing student classroom behaviors is essential for analyzing teacher\u2013student interactions and enabling intelligent educational assessment. Although deep learning offers promising solutions, existing methods often perform poorly in complex classroom environments due to occlusions and subtle, overlapping actions. To address these issues, this article proposes a robust and efficient method for behavior recognition by enhancing the You Only Look Once version 8 (YOLOv8) architecture with a Multi-Head Self-Attention (MHSA) module, termed YOLOv8-MHSA. The integration of MHSA allows the model to capture contextual relationships between distant spatial features, which is critical for distinguishing similar behaviors. For a comprehensive evaluation, we also implement a model with Coordinate Attention (CA). Experimental results on a standard dataset demonstrate the superiority of our YOLOv8-MHSA model, which achieves a precision of 0.86, recall of 0.807, mAP50 of 0.855, and mAP50-95 of 0.677, delivering competitive performance compared to the state-of-the-art SBD-Net. These findings validate that explicit contextual modeling via self-attention significantly boosts performance in fine-grained behavior recognition. Consequently, this research has direct potential applications in providing automated, data-driven tools for teacher training, classroom quality assessment, and, ultimately, supporting the development of personalized education systems.<\/jats:p>","DOI":"10.3390\/info16110934","type":"journal-article","created":{"date-parts":[[2025,10,30]],"date-time":"2025-10-30T03:44:39Z","timestamp":1761795879000},"page":"934","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Student Classroom Behavior Recognition Based on YOLOv8 and Attention Mechanism"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8006-747X","authenticated-orcid":false,"given":"Jingpu","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Computer and Data Science, Henan University of Urban Construction, Pingdingshan 467036, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6279-0338","authenticated-orcid":false,"given":"Lizheng","family":"Guo","sequence":"additional","affiliation":[{"name":"School of Computer and Data Science, Henan University of Urban Construction, Pingdingshan 467036, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xuyang","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer and Data Science, Henan University of Urban Construction, Pingdingshan 467036, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,10,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Wang, Z., Yao, J., Zeng, C., Li, L., and Tan, C. (2023). Students\u2019 Classroom Behavior Detection System Incorporating Deformable DETR with Swin Transformer and Light-Weight Feature Pyramid Network. Systems, 11.","DOI":"10.3390\/systems11070372"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"49","DOI":"10.1038\/s41586-024-07146-0","article-title":"Artificial Intelligence and Illusions of Understanding in Scientific Research","volume":"627","author":"Messeri","year":"2024","journal-title":"Nature"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"2553","DOI":"10.18280\/ts.400618","article-title":"Analyzing and Optimizing Virtual Reality Classroom Scenarios: A Deep Learning Approach","volume":"40","author":"Jiang","year":"2023","journal-title":"Trait. Signal"},{"key":"ref_4","unstructured":"Yang, F. (2023). SCB-dataset: A dataset for detecting student classroom behavior. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"164","DOI":"10.1016\/j.procs.2024.03.206","article-title":"Attention-Based AdaptSepCX Network for Effective Student Action Recognition in Online Learning","volume":"233","author":"Dey","year":"2024","journal-title":"Procedia Comput. Sci."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"120","DOI":"10.1097\/01.NEP.0000000000001086","article-title":"Evidence-Based Classroom Observation Technique: An Interdisciplinary, Structured Approach to Classroom Observation","volume":"45","author":"Perkins","year":"2024","journal-title":"Nurs. Educ. Perspect."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"32181","DOI":"10.1109\/ACCESS.2024.3368855","article-title":"Telepresence Observation for Kindergarten Classroom Rating: A Pilot Study","volume":"12","author":"Lu","year":"2024","journal-title":"IEEE Access"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"104726","DOI":"10.1016\/j.imavis.2023.104726","article-title":"Student Behavior Recognition for Interaction Detection in the Classroom Environment","volume":"136","author":"Li","year":"2023","journal-title":"Image Vis. Comput."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1007\/s11092-017-9261-5","article-title":"Resistance to Classroom Observation in the Context of Teacher Evaluation: Teachers\u2019 and Department Heads\u2019 Experiences and Perspectives","volume":"30","author":"Silva","year":"2018","journal-title":"Educ. Assess. Eval. Account."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1007\/s10639-024-12734-8","article-title":"Deciphering the impact of machine learning on education: Insights from a bibliometric analysis using bibliometrix R-package","volume":"29","author":"Zhong","year":"2024","journal-title":"Educ. Inf. Technol."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"5990","DOI":"10.1109\/TPAMI.2025.3559891","article-title":"MB-TaylorFormer V2: Improved Multi-Branch Linear Transformer Expanded by Taylor Formula for Image Restoration","volume":"47","author":"Jin","year":"2025","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"4541","DOI":"10.1007\/s11263-024-02056-0","article-title":"GridFormer: Residual Dense Transformer with Grid Structure for Image Restoration in Adverse Weather Conditions","volume":"132","author":"Wang","year":"2024","journal-title":"Int. J. Comput. Vis."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1287","DOI":"10.1109\/TPAMI.2022.3148707","article-title":"Enhanced Spatio-Temporal Interaction Learning for Video Deraining: A Faster and Better Framework","volume":"45","author":"Zhang","year":"2023","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"7419","DOI":"10.1109\/TIP.2021.3104166","article-title":"Deep Dense Multi-Scale Network for Snow Removal Using Semantic and Depth Priors","volume":"30","author":"Zhang","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"111628","DOI":"10.1016\/j.patcog.2025.111628","article-title":"LLDiffusion: Learning Degradation Representations in Diffusion Models for Low-Light Image Enhancement","volume":"166","author":"Wang","year":"2025","journal-title":"Pattern Recogn."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1454","DOI":"10.1037\/edu0000644","article-title":"Cognitive dimensions of learning in children with problems in attention, learning, and memory","volume":"113","author":"Holmes","year":"2021","journal-title":"J. Educ. Psychol."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Wang, Z., Wang, M., Zeng, C., and Li, L. (2024). SBD-Net: Incorporating Multi-Level Features for an Efficient Detection Network of Student Behavior in Smart Classrooms. Appl. Sci., 14.","DOI":"10.3390\/app14188357"},{"key":"ref_18","first-page":"1407","article-title":"A Systematic Review for the Fatigue Driving Behavior Recognition Method","volume":"46","author":"Hou","year":"2024","journal-title":"J. Intell. Fuzzy Syst."},{"key":"ref_19","first-page":"122","article-title":"Revolutionizing Political Education in Pakistan: An AI-Integrated Approach","volume":"1","author":"Saqlain","year":"2023","journal-title":"Educ. Sci. Manag."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"331","DOI":"10.1089\/tmj.2022.0037","article-title":"Use of Technologies in the Therapy of Social Cognition Deficits in Neurological and Mental Diseases: A Systematic Review","volume":"29","author":"Lohaus","year":"2023","journal-title":"Telemed. E-Health"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Tang, L., Xie, T., Yang, Y., and Wang, H. (2022). Classroom Behavior Detection Based on Improved YOLOv5 Algorithm Combining Multi-Scale Feature Fusion and Attention Mechanism. Appl. Sci., 12.","DOI":"10.3390\/app12136790"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"19091","DOI":"10.1007\/s11042-022-14100-7","article-title":"Student Behavior Recognition Based on Multitask Learning","volume":"82","author":"Mo","year":"2023","journal-title":"Multimed. Tools Appl."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zong, L., and Fang, J. (2024). Deep Visual Computing of Behavioral Characteristics in Complex Scenarios and Embedded Object Recognition Applications. Sensors, 24.","DOI":"10.3390\/s24144582"},{"key":"ref_24","first-page":"9903342","article-title":"Identifying and Monitoring Students\u2019 Classroom Learning Behavior Based on Multisource Information","volume":"2022","author":"Sun","year":"2022","journal-title":"Mob. Inf. Syst."},{"key":"ref_25","first-page":"52","article-title":"Student Engagement Detection Using Emotion Analysis, Eye Tracking and Head Movement with Machine Learning","volume":"Volume 1720","author":"Reis","year":"2022","journal-title":"Technology and Innovation in Learning, Teaching and Education"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Delgado, K., Origgi, J.M., Hasanpoor, T., Yu, H., Allessio, D., Arroyo, I., Lee, W., Betke, M., Woolf, B., and Bargal, S.A. (2021, January 11\u201317). Student Engagement Dataset. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV) Workshops, Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00405"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Gu, C., Sun, C., Ross, D.A., Vondrick, C., Pantofaru, C., Li, Y., Vijayanarasimhan, S., Toderici, G., Ricco, S., and Sukthankar, R. (2018, January 18\u201323). Ava: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00633"},{"key":"ref_28","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., and He, K. (November, January 27). Slowfast Networks for Video Recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Sigurdsson, G.A., Varol, G., Wang, X., Farhadi, A., Laptev, I., and Gupta, A. (2016, January 11\u201314). Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding. Proceedings of the Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands. Proceedings, Part I 14.","DOI":"10.1007\/978-3-319-46448-0_31"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Carreira, J., and Zisserman, A. (2017, January 21\u201326). Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.502"},{"key":"ref_31","first-page":"133","article-title":"A New Feature Fusion Network for Student Behavior Recognition in Education","volume":"24","author":"Jisi","year":"2021","journal-title":"J. Appl. Sci. Eng."},{"key":"ref_32","first-page":"8","article-title":"A Recognition Method of Learning Behaviour in English Online Classroom Based on Feature Data Mining","volume":"15","author":"Shi","year":"2023","journal-title":"Int. J. Reason.-Based Intell. Syst."},{"key":"ref_33","first-page":"674","article-title":"Student Behavior Analysis using YOLOv5 and OpenPose in Smart Classroom Environment","volume":"2024","author":"Li","year":"2025","journal-title":"AMIA Annu. Symp. Proc."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Sheng, X., Li, S., and Chan, S. (2025). Real-time classroom student behavior detection based on improved YOLOv8s. Sci. Rep., 15.","DOI":"10.1038\/s41598-025-99243-x"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"2907","DOI":"10.1007\/s11042-020-09741-5","article-title":"Surveillance Video Analysis for Student Action Recognition and Localization inside Computer Laboratories of a Smart Campus","volume":"80","author":"Rashmi","year":"2021","journal-title":"Multimed. Tools Appl."},{"key":"ref_36","first-page":"757","article-title":"Student Activities Detection of SUST Using YOLOv3 on Deep Learning","volume":"8","author":"Ali","year":"2020","journal-title":"Indones. J. Electr. Eng. Inform. IJEEI"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1016\/j.cag.2022.11.009","article-title":"Hand-Raising Gesture Detection in Classroom with Spatial Context Augmentation and Dilated Convolution","volume":"110","author":"Zhang","year":"2023","journal-title":"Comput. Graph."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Wang, Z., Li, L., Zeng, C., and Yao, J. (2023). Student Learning Behavior Recognition Incorporating Data Augmentation with Learning Feature Representation in Smart Classrooms. Sensors, 23.","DOI":"10.3390\/s23198190"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Chen, H., Zhou, G., and Jiang, H. (2023). Student Behavior Detection in the Classroom Based on Improved YOLOv8. Sensors, 23.","DOI":"10.3390\/s23208385"},{"key":"ref_40","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"219","DOI":"10.1080\/00461520.2014.965823","article-title":"The ICAP framework: Linking cognitive engagement to active learning outcomes","volume":"49","author":"Chi","year":"2014","journal-title":"Educ. Psychol."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/11\/934\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,30]],"date-time":"2025-10-30T04:03:31Z","timestamp":1761797011000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/11\/934"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,27]]},"references-count":41,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2025,11]]}},"alternative-id":["info16110934"],"URL":"https:\/\/doi.org\/10.3390\/info16110934","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,27]]}}}