{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,14]],"date-time":"2026-05-14T00:04:41Z","timestamp":1778717081689,"version":"3.51.4"},"reference-count":33,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2025,11,3]],"date-time":"2025-11-03T00:00:00Z","timestamp":1762128000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Scientific Research Program of the Hubei Provincial Department of Education","award":["B2023417"],"award-info":[{"award-number":["B2023417"]}]},{"name":"Hubei Province First-Class Undergraduate Course Construction Project"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>The ongoing integration of information technology in education has rendered the monitoring of student behavior in smart classrooms essential for improving teaching quality and student engagement. Classroom environments frequently provide many problems, such as heterogeneous student behaviors, significant obstructions, loss of intricate details, and complications in recognizing diminutive targets. These limitations lead to current approaches remaining inadequate in accuracy and stability. This paper enhances YOLOv11 with the following improvements: developed the CSP-PMSA module to enhance contextual modeling in complex backgrounds, developed a scale-aware head (SAH) to improve the perception and localization of small targets via channel unification and scale adaptation, and introduced a Multi-Head Self-Attention (MHSA) mechanism to model global dependencies and positional bias across various subspaces, thereby enhancing the discrimination of visually analogous behaviors. The experimental findings indicate that in intricate classroom settings, the model attains mAP@50 and mAP@50\u201395 scores of 91.6% and 75.7%, respectively. This indicates enhancements of 2.7% and 2.6% compared to YOLOv11, and 4.6% and 3.6% relative to DETR, demonstrating remarkable detection precision and dependability. Additionally, the model was implemented on the Jetson Orin Nano platform, confirming its viability for real-time detection on edge devices and offering substantial assistance for practical implementations in smart classrooms.<\/jats:p>","DOI":"10.3390\/info16110949","type":"journal-article","created":{"date-parts":[[2025,11,3]],"date-time":"2025-11-03T19:30:27Z","timestamp":1762198227000},"page":"949","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Deep Learning for Student Behavior Detection in Smart Classroom Environments"],"prefix":"10.3390","volume":"16","author":[{"given":"Jue","family":"Wang","sequence":"first","affiliation":[{"name":"School of Design Engineering, Wuhan Qingchuan University, Wuhan 430204, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuchen","family":"Sun","sequence":"additional","affiliation":[{"name":"School of Teaching, Learning and Curriculum Studies, Kent State University, 800 E Summit St, Kent, OH 44240, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shasha","family":"Tian","sequence":"additional","affiliation":[{"name":"School of Computer Science, South-Central Minzu University, Wuhan 430074, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,11,3]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1129","DOI":"10.1007\/s40593-024-00422-0","article-title":"Artificial intelligence for enhancing special education for K-12: A decade of trends, themes, and global insights (2013\u20132023)","volume":"35","author":"Yang","year":"2024","journal-title":"Int. J. Artif. Intell. Educ."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1007\/s44196-024-00572-y","article-title":"Real-time classroom behavior analysis for enhanced engineering education: An AI-assisted approach","volume":"17","author":"Hu","year":"2024","journal-title":"Int. J. Comput. Intell. Syst."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Liu, Q., Jiang, X., and Jiang, R. (2025). Classroom behavior recognition using computer vision: A systematic review. Sensors, 25.","DOI":"10.3390\/s25020373"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"49767","DOI":"10.1109\/ACCESS.2025.3550921","article-title":"Classroom student behavior recognition using an intelligent sensing framework","volume":"13","author":"Zhao","year":"2025","journal-title":"IEEE Access"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"12861","DOI":"10.1007\/s11227-022-04402-w","article-title":"An improved method of identifying learner\u2019s behaviors based on deep learning","volume":"78","author":"Liu","year":"2022","journal-title":"J. Supercomput."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"110114","DOI":"10.1016\/j.asoc.2023.110114","article-title":"Evolutionary machine learning builds smart education big data platform: Data-driven higher education","volume":"136","author":"Zheng","year":"2023","journal-title":"Appl. Soft Comput."},{"key":"ref_7","first-page":"6336773","article-title":"A convolutional neural network (CNN) based approach for the recognition and evaluation of classroom teaching behavior","volume":"2021","author":"Li","year":"2021","journal-title":"Sci. Program."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Feng, C., Luo, Z., Kong, D., Ding, Y., and Liu, J. (2025). IMRMB-Net: A lightweight student behavior recognition model for complex classroom scenarios. PLoS ONE, 20.","DOI":"10.1371\/journal.pone.0318817"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Han, L., Ma, X., Dai, M., and Bai, L. (2025). A WAD-YOLOv8-based method for classroom student behavior detection. Sci. Rep., 15.","DOI":"10.1038\/s41598-025-87661-w"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Chen, H., Zhou, G., and Jiang, H. (2023). Student behavior detection in the classroom based on improved YOLOv8. Sensors, 23.","DOI":"10.3390\/s23208385"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_13","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015). Faster R-CNN: Towards real-time object detection with region proposal networks. Adv. Neural Inf. Process. Syst., 28."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A.C. (2016, January 8\u201316). SSD: Single shot multibox detector. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_16","unstructured":"Redmon, J., and Farhadi, A. (2018). YOLOv3: An incremental improvement. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201322). Path aggregation network for instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"43","DOI":"10.11648\/j.ijsedu.20150305.11","article-title":"The instructional process: A review of Flanders\u2019 interaction analysis in a classroom setting","volume":"3","author":"Amatari","year":"2015","journal-title":"Int. J. Second. Educ."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wang, D., Han, H., and Liu, H. (2019, January 6\u20138). Analysis of instructional interaction behaviors based on OOTIAS in smart learning environment. Proceedings of the 2019 Eighth International Conference on Educational Innovation through Technology (EITT), Biloxi, MS, USA.","DOI":"10.1109\/EITT.2019.00036"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"25","DOI":"10.28991\/esj-2021-01254","article-title":"Human action recognition in videos using convolution long short-term memory network with spatio-temporal networks","volume":"5","author":"Sarabu","year":"2021","journal-title":"Emerg. Sci. J."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"8699","DOI":"10.1109\/TMM.2023.3239751","article-title":"Skeleton-based action recognition through contrasting two-stream spatial-temporal networks","volume":"25","author":"Pang","year":"2023","journal-title":"IEEE Trans. Multimed."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"42117","DOI":"10.1007\/s11042-021-11220-4","article-title":"Deep convolutional neural model for human activities recognition in a sequence of video by combining multiple CNN streams","volume":"81","author":"Varshney","year":"2022","journal-title":"Multimed. Tools Appl."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wang, Z., Yao, J., Zeng, C., Li, L., and Tan, C. (2023). Students\u2019 classroom behavior detection system incorporating deformable DETR with Swin transformer and light-weight feature pyramid network. Systems, 11.","DOI":"10.3390\/systems11070372"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Peng, S., Zhang, X., Zhou, L., and Wang, P. (2025). YOLO-CBD: Classroom behavior detection method based on behavior feature extraction and aggregation. Sensors, 25.","DOI":"10.3390\/s25103073"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"140","DOI":"10.1007\/s11554-024-01515-8","article-title":"CSB-YOLO: A rapid and efficient real-time algorithm for classroom student behavior detection","volume":"21","author":"Zhu","year":"2024","journal-title":"J. Real-Time Image Process."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Chen, H., and Guan, J. (2022). Teacher\u2013student behavior recognition in classroom teaching based on improved YOLO-v4 and Internet of Things technology. Electronics, 11.","DOI":"10.3390\/electronics11233998"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Wu, Z., Chen, X., Dai, L., Li, Z., Zong, X., and Liu, T. (2020, January 26\u201328). Classroom behavior recognition based on improved YOLOv3. Proceedings of the 2020 International Conference on Artificial Intelligence and Education (ICAIE), Hangzhou, China.","DOI":"10.1109\/ICAIE50891.2020.00029"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Ouyang, D., He, S., Zhang, G., Luo, M., Guo, H., Zhan, J., and Huang, Z. (2023, January 4\u201310). Efficient multi-scale attention module with cross-spatial learning. Proceedings of the ICASSP 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece.","DOI":"10.1109\/ICASSP49357.2023.10096516"},{"key":"ref_30","unstructured":"Yang, L., Zhang, R.-Y., Li, L., and Xie, X. (2021, January 18\u201324). SimAM: A simple, parameter-free attention module for convolutional neural networks. Proceedings of the International Conference on Machine Learning (ICML), Virtual Event."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"121352","DOI":"10.1016\/j.eswa.2023.121352","article-title":"Large separable kernel attention: Rethinking the large kernel attention design in CNN","volume":"236","author":"Lau","year":"2024","journal-title":"Expert Syst. Appl."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Cai, X., Lai, Q., Wang, Y., Wang, W., Sun, Z., and Yao, Y. (2024, January 17\u201321). Poly kernel inception network for remote sensing detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02617"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Srinivas, A., Lin, T.-Y., Parmar, N., Shlens, J., Abbeel, P., and Vaswani, A. (2021, January 19\u201325). Bottleneck transformers for visual recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01625"}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/11\/949\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,3]],"date-time":"2025-11-03T19:48:26Z","timestamp":1762199306000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/11\/949"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,3]]},"references-count":33,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2025,11]]}},"alternative-id":["info16110949"],"URL":"https:\/\/doi.org\/10.3390\/info16110949","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,3]]}}}