{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T04:23:42Z","timestamp":1783571022496,"version":"3.55.0"},"reference-count":49,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T00:00:00Z","timestamp":1740096000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Journal of Ambient Intelligence and Smart Environments"],"published-print":{"date-parts":[[2025,8]]},"abstract":"<jats:p>Virtual education (online education or e-learning) is a form of education where the primary mode of instruction is through digital platforms and the Internet. This approach offers flexibility and accessibility, making it attractive to many students. Many institutes also offer virtual professional courses for business and working professionals. However, ensuring the reachability of courses and evaluating students\u2019 attentiveness presents significant challenges for educators teaching virtually. Various research works have been proposed to evaluate students\u2019 attentiveness using facial landmarks, facial expressions, eye movements, gestures, postures, etc. However, no method has been proposed for real-time analysis and evaluation. This paper introduces a multi-modal student attentiveness detection (MMSAD) model designed to analyze and evaluate real-time class videos using two modalities: facial expressions and landmarks. Using a lightweight deep learning model, the model analyzes students\u2019 emotions from facial expressions and identifies when a person is speaking during an online class by examining lip movements from facial landmarks. The model evaluates students\u2019 emotions using five benchmark datasets, achieving accuracy rates of 99.05% on extended Cohn-Kanade (CK+), 87.5% on RAF-DB, 78.12% on Facial Emotion Recognition-2013 (FER-2013), 98.50% on JAFFE, and 88.01% on KDEF. The model identifies individuals speaking during the class using real-time class videos. The results from these modalities are used to predict attentiveness, categorizing students as either attentive or inattentive.<\/jats:p>","DOI":"10.1177\/18761364251315239","type":"journal-article","created":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T04:50:50Z","timestamp":1740113450000},"page":"326-348","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":3,"title":["MMSAD\u2014A multi-modal student attentiveness detection in smart education using facial features and landmarks"],"prefix":"10.1177","volume":"17","author":[{"given":"Ruchi","family":"Singh","sequence":"first","affiliation":[{"name":"Computer Science and Engineering, National Institute of Technology Silchar, Assam, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1450-9097","authenticated-orcid":false,"given":"Ramanujam","family":"E","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, National Institute of Technology Silchar, Assam, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Naresh Babu","family":"M","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, National Institute of Technology Silchar, Assam, India"},{"name":"Computer Science and Engineering, Indian Institute of Information Technology, Design and Manufacturing Kurnool, Andhra Pradesh, India"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"179","published-online":{"date-parts":[[2025,2,21]]},"reference":[{"key":"e_1_3_3_2_1","doi-asserted-by":"crossref","unstructured":"Abdellaoui B Moumen A El Idrissi YEB et al. (2020) Face detection to recognize students\u2019 emotion and their engagement: A systematic review. In: 2020 IEEE 2nd international conference on electronics control optimization and computer science (ICECOCS) Kenitra Morocco pp.1\u20136. IEEE.","DOI":"10.1109\/ICECOCS50124.2020.9314600"},{"key":"e_1_3_3_3_1","doi-asserted-by":"crossref","unstructured":"Akinrinmade AA Adetiba E Badejo JA et al. (2023) An active speaker detection method in videos using standard deviations of color histogram. In: 2023 International conference on science engineering and business for sustainable development goals (SEB-SDG) \u00a0Omu-Aran Nigeria vol. 1 pp.1\u20136.","DOI":"10.1109\/SEB-SDG57117.2023.10124488"},{"key":"e_1_3_3_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3325407"},{"key":"e_1_3_3_5_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0258788"},{"key":"e_1_3_3_6_1","doi-asserted-by":"crossref","unstructured":"Azizi FN Kurniawardhani A Paputungan IV (2022) Facial expression image based emotion detection using convolutional neural network. In: 2022 IEEE 20th student conference on research and development (SCOReD) Bangi Malaysia pp.157\u2013162.","DOI":"10.1109\/SCOReD57082.2022.9974104"},{"key":"e_1_3_3_7_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.compeleceng.2021.107277"},{"key":"e_1_3_3_8_1","doi-asserted-by":"publisher","DOI":"10.3758\/BRM.40.1.109"},{"key":"e_1_3_3_9_1","doi-asserted-by":"publisher","DOI":"10.3390\/app10010314"},{"key":"e_1_3_3_10_1","doi-asserted-by":"publisher","DOI":"10.1186\/s40561-018-0080-z"},{"key":"e_1_3_3_11_1","doi-asserted-by":"publisher","DOI":"10.1080\/02522667.2020.1809126"},{"key":"e_1_3_3_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10639-023-11814-5"},{"key":"e_1_3_3_13_1","doi-asserted-by":"crossref","unstructured":"Goodfellow ID Erhan J Carrier PL et al. (2013) Challenges in representation learning: A report on three machine learning contests. In: Lee M Hirose A Hou Z-G and Kil HM (eds) Neural Information Processing. Berlin: Springer pp.117\u2013124.","DOI":"10.1007\/978-3-642-42051-1_16"},{"key":"e_1_3_3_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-023-14392-3"},{"key":"e_1_3_3_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-022-13558-9"},{"key":"e_1_3_3_16_1","doi-asserted-by":"crossref","unstructured":"Haider F Campbell N Luz S (2016) Active speaker detection in human machine multiparty dialogue using visual prosody information. In: 2016 IEEE global conference on signal and information processing (GlobalSIP) Washington DC USA pp.1207\u20131211. IEEE.","DOI":"10.1109\/GlobalSIP.2016.7906033"},{"key":"e_1_3_3_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2021.09.115"},{"key":"e_1_3_3_18_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.asoc.2019.105724"},{"key":"e_1_3_3_19_1","first-page":"1755","article-title":"DLIB-ML: A machine learning toolkit","volume":"10","author":"King DE","year":"2009","unstructured":"King DE (2009) DLIB-ML: A machine learning toolkit. Journal of Machine Learning Research 10: 1755\u20131758.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_3_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10639-022-11370-4"},{"key":"e_1_3_3_21_1","doi-asserted-by":"publisher","DOI":"10.3390\/app13116409"},{"key":"e_1_3_3_22_1","doi-asserted-by":"crossref","unstructured":"Li S Deng W Du J (2017) Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild. In: 2017 IEEE conference on computer vision and pattern recognition (CVPR) \u00a0Honolulu HI USA pp.2584\u20132593.","DOI":"10.1109\/CVPR.2017.277"},{"key":"e_1_3_3_23_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10489-020-02139-8"},{"key":"e_1_3_3_24_1","doi-asserted-by":"crossref","unstructured":"Liu C Jiang W Wang M et al. (2020) Group level audio-video emotion recognition using hybrid networks. In: Proceedings of the 2020 International conference on multimodal interaction (ICMI\u201920) New York NY USA pp.807\u2013812. Association for Computing Machinery.","DOI":"10.1145\/3382507.3417968"},{"key":"e_1_3_3_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-018-6017-2"},{"key":"e_1_3_3_26_1","doi-asserted-by":"crossref","unstructured":"Liu Y Zeng J Shan S et al. (2018b) Multi-channel pose-aware convolution neural networks for multi-view facial expression recognition. In: 2018 13th IEEE international conference on automatic face & gesture recognition (FG 2018) Xi'an China pp.458\u2013465. IEEE.","DOI":"10.1109\/FG.2018.00074"},{"key":"e_1_3_3_27_1","doi-asserted-by":"crossref","unstructured":"Lucey P Cohn JF Kanade T et al. (2010) The extended Cohn-Kanade dataset (CK+): A complete dataset for action unit and emotion-specified expression. In: 2010 IEEE computer society conference on computer vision and pattern recognition \u2013 workshops San Francisco CA USA pp.94\u2013101.","DOI":"10.1109\/CVPRW.2010.5543262"},{"key":"e_1_3_3_28_1","doi-asserted-by":"crossref","unstructured":"Lyons M Akamatsu S Kamachi M et al. (1998) Coding facial expressions with Gabor wavelets. In: Proceedings third IEEE international conference on automatic face and gesture recognition Nara Japan pp.200\u2013205.","DOI":"10.1109\/AFGR.1998.670949"},{"key":"e_1_3_3_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/s13278-023-01181-x"},{"key":"e_1_3_3_30_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2016.07.233"},{"key":"e_1_3_3_31_1","doi-asserted-by":"crossref","unstructured":"Mukhopadhyay M Pal S Nayyar A et al. (2020) Facial emotion detection to assess learner\u2019s state of mind in an online learning system. In: Proceedings of the 2020 5th international conference on intelligent information technology (ICIIT\u201920) New York NY USA pp.107\u2013115. Association for Computing Machinery.","DOI":"10.1145\/3385209.3385231"},{"key":"e_1_3_3_32_1","doi-asserted-by":"crossref","unstructured":"Murshed M Dewan MAA Lin F et al. (2019) Engagement detection in e-learning environments using convolutional neural networks. In: 2019 IEEE international conference on dependable autonomic and secure computing international conference on pervasive intelligence and computing international conference on cloud and big data computing international conference on cyber science and technology congress (DASC\/PiCom\/CBDCom\/CyberSciTech) Fukuoka Japan pp.80\u201386.","DOI":"10.1109\/DASC\/PiCom\/CBDCom\/CyberSciTech.2019.00028"},{"key":"e_1_3_3_33_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijedro.2020.100011"},{"key":"e_1_3_3_34_1","doi-asserted-by":"publisher","DOI":"10.3390\/s20041087"},{"key":"e_1_3_3_35_1","doi-asserted-by":"crossref","unstructured":"Sari M Moussaoui A Hadid A (2021) A simple yet effective convolutional neural network model to classify facial expressions. In: Modelling and implementation of complex systems: Proceedings of the 6th international symposium MISC 2020 Batna Algeria October 24\u201326 2020 pp.188\u2013202. Springer.","DOI":"10.1007\/978-3-030-58861-8_14"},{"key":"e_1_3_3_36_1","doi-asserted-by":"crossref","unstructured":"Shah NA Meenakshi K Agarwal A et al. (2021) Assessment of student attentiveness to e-learning by monitoring behavioural elements. In: 2021 International conference on computer communication and informatics (ICCCI) Coimbatore India pp.1\u20137.","DOI":"10.1109\/ICCCI50826.2021.9402283"},{"key":"e_1_3_3_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-021-00854-x"},{"key":"e_1_3_3_38_1","doi-asserted-by":"crossref","unstructured":"Tabassum T Allen AA De P (2020) Non-intrusive identification of student attentiveness and finding their correlation with detectable facial emotions. In: Proceedings of the 2020 ACM southeast conference Tampa FL USA ACM SE\u201920. ACM.","DOI":"10.1145\/3374135.3385263"},{"key":"e_1_3_3_39_1","doi-asserted-by":"crossref","unstructured":"Viola P Jones M (2001) Rapid object detection using a boosted cascade of simple features. In: Proceedings of the 2001 IEEE computer society conference on computer vision and pattern recognition. CVPR 2001 Kauai HI USA vol. 1 pp.I\u2013I. IEEE.","DOI":"10.1109\/CVPR.2001.990517"},{"key":"e_1_3_3_40_1","doi-asserted-by":"publisher","DOI":"10.3389\/feduc.2022.1074435"},{"key":"e_1_3_3_41_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.107837"},{"key":"e_1_3_3_42_1","doi-asserted-by":"publisher","DOI":"10.1080\/10447318.2021.1938389"},{"key":"e_1_3_3_43_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2020.107694"},{"key":"e_1_3_3_44_1","doi-asserted-by":"crossref","unstructured":"Wu Y Zhang L Chen G et al. (2021) Unconstrained facial expression recognition based on cascade decision and Gabor filters. In: 2020 25th International conference on pattern recognition (ICPR) Milan Italy pp.3336\u20133341. IEEE.","DOI":"10.1109\/ICPR48806.2021.9411983"},{"key":"e_1_3_3_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-023-01153-3"},{"key":"e_1_3_3_46_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10639-023-12058-z"},{"key":"e_1_3_3_47_1","doi-asserted-by":"crossref","unstructured":"Zakka BE Vadapalli H (2020) Estimating student learning affect using facial emotions. In: 2020 2nd international multidisciplinary information technology and engineering conference (IMITEC) \u00a0Kimberley South Africa pp.1\u20136","DOI":"10.1109\/IMITEC50163.2020.9334075"},{"key":"e_1_3_3_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCDS.2020.3034807"},{"key":"e_1_3_3_49_1","doi-asserted-by":"publisher","DOI":"10.1177\/0735633119825575"},{"key":"e_1_3_3_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11760-024-03020-8"}],"container-title":["Journal of Ambient Intelligence and Smart Environments"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/18761364251315239","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/18761364251315239","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/18761364251315239","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T09:21:15Z","timestamp":1777368075000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/18761364251315239"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,21]]},"references-count":49,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,8]]}},"alternative-id":["10.1177\/18761364251315239"],"URL":"https:\/\/doi.org\/10.1177\/18761364251315239","relation":{},"ISSN":["1876-1364","1876-1372"],"issn-type":[{"value":"1876-1364","type":"print"},{"value":"1876-1372","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,21]]}}}