{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T16:38:53Z","timestamp":1779295133085,"version":"3.51.4"},"reference-count":17,"publisher":"Wiley","issue":"1","license":[{"start":{"date-parts":[[2026,1,15]],"date-time":"2026-01-15T00:00:00Z","timestamp":1768435200000},"content-version":"vor","delay-in-days":14,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T00:00:00Z","timestamp":1767225600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Internet Technology Letters"],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>In the English teaching environment driven by the Internet of Things, real\u2010time assessment of students' emotional states based on video data is particularly crucial for understanding student engagement and improving teaching quality. Existing server\u2010based deep networks are limited by long\u2010distance video data transmission, which seriously restricts the real\u2010time performance of sentiment analysis. Moreover, simple video data cannot guarantee robustness in complex scenes. To address these issues, this paper proposes a multimodal fusion emotion recognition framework based on the edge\u2010cloud collaboration mechanism. Firstly, on the edge node, we exploit two complementary modalities of data: video sequences and facial landmark sequences, and design a lightweight dual\u2010stream neural network based on the 3D MobileNetV3 and graph convolutional network to efficiently extract multimodal features. On the server, we adopt the Transformer\u2010based cross fusion mechanism to implement multimodal fusion and emotion evaluation. The edge side is responsible for real\u2010time preprocessing and primary feature extraction. In our proposed framework, the server is responsible for aggregating feature data from multiple edge nodes. The experimental results indicate that the proposed framework can achieve high\u2010precision student engagement assessment with low latency.<\/jats:p>","DOI":"10.1002\/itl2.70223","type":"journal-article","created":{"date-parts":[[2026,1,15]],"date-time":"2026-01-15T13:50:12Z","timestamp":1768485012000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["A Multimodal Affective Computing Framework for Real\u2010Time Student Engagement Assessment in\n                    <scp>IoT<\/scp>\n                    \u2010Enabled English Classrooms: An Edge\u2010Cloud Collaborative Approach"],"prefix":"10.1002","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-0049-7819","authenticated-orcid":false,"given":"Peirong","family":"He","sequence":"first","affiliation":[{"name":"Hunan Vocational College of Commerce  Changsha Hunan China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,1,15]]},"reference":[{"key":"e_1_2_8_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.system.2018.03.015"},{"key":"e_1_2_8_3_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00500-022-07156-y"},{"key":"e_1_2_8_4_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2024.126236"},{"key":"e_1_2_8_5_1","doi-asserted-by":"crossref","unstructured":"U.Chindiyababy P.Kakkar J.Vedula J.Yunus A.Umidbek andS.Sharma \u201cDeep Learning\u2010Based Facial Emotion Recognition for Advanced Human\u2010Computer Interaction \u201dIEEE 2025 1247\u20131252.","DOI":"10.1109\/CE2CT64011.2025.10941356"},{"key":"e_1_2_8_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2025.107901"},{"key":"e_1_2_8_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2025.3553413"},{"key":"e_1_2_8_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCE.2025.3568882"},{"key":"e_1_2_8_9_1","doi-asserted-by":"publisher","DOI":"10.3390\/s25165066"},{"key":"e_1_2_8_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2025.3555510"},{"key":"e_1_2_8_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2020.03.005"},{"key":"e_1_2_8_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2984368"},{"key":"e_1_2_8_13_1","doi-asserted-by":"crossref","unstructured":"A.Zadeh P. P.Liang S.Poria P.Vij E.Cambria andL. P.Morency \u201cMulti\u2010Attention Recurrent Network for Human Communication Comprehension \u201d2018 5642\u20135649.","DOI":"10.1609\/aaai.v32i1.12024"},{"key":"e_1_2_8_14_1","doi-asserted-by":"crossref","unstructured":"L.Zhang S.Walter X.Ma et al. \u201c\u201cBioVid Emo DB\u201d: A Multimodal Database for Emotion Analyses Validated by Subjective Ratings \u201dIEEE 2016 1\u20136.","DOI":"10.1109\/SSCI.2016.7849931"},{"key":"e_1_2_8_15_1","unstructured":"M.SinghandY.Fang \u201cEmotion Recognition in Audio and Video Using Deep Neural Networks \u201darXiv Preprint 2020 arXiv:2006.08129."},{"key":"e_1_2_8_16_1","doi-asserted-by":"crossref","unstructured":"K.Rupauliha A.Goyal A.Saini A.Shukla andS.Swaminathan \u201cMultimodal Emotion Recognition in Polish (Student Consortium) \u201dIEEE 2020 307\u2013311.","DOI":"10.1109\/BigMM50055.2020.00054"},{"key":"e_1_2_8_17_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i02.5492"},{"key":"e_1_2_8_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2023.3275156"}],"container-title":["Internet Technology Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/itl2.70223","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1002\/itl2.70223","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/itl2.70223","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,23]],"date-time":"2026-01-23T03:39:33Z","timestamp":1769139573000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/itl2.70223"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1]]},"references-count":17,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["10.1002\/itl2.70223"],"URL":"https:\/\/doi.org\/10.1002\/itl2.70223","archive":["Portico"],"relation":{},"ISSN":["2476-1508","2476-1508"],"issn-type":[{"value":"2476-1508","type":"print"},{"value":"2476-1508","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1]]},"assertion":[{"value":"2025-10-30","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-05","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70223"}}