{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,18]],"date-time":"2026-03-18T21:51:40Z","timestamp":1773870700339,"version":"3.50.1"},"reference-count":43,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,6,27]],"date-time":"2025-06-27T00:00:00Z","timestamp":1750982400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Artif. Intell."],"abstract":"<jats:p>This study aims to assess participant engagement in multiparty conversations using video and audio data. For this task, the interaction among numerous data streams, such as video and audio from multiple participants, should be modeled effectively, considering the redundancy of video and audio across frames. To efficiently model participant interactions while accounting for such redundancy, a previous study proposed inputting participant feature sequences into global token-based transformers, which constrain attention across feature sequences to pass through only a small set of internal units, allowing the model to focus on key information. However, this approach still faces the challenge of redundancy in participant-feature estimation based on standard cross-attention transformers, which can connect all frames across different modalities. To address this, we propose a joint model for interactions among all data streams using global token-based transformers, without distinguishing between cross-modal and cross-participant interactions. Experiments on the RoomReader corpus confirm that the proposed model outperforms previous models, achieving accuracy ranging from 0.720 to 0.763, weighted F1 scores from 0.733 to 0.771, and macro F1 scores from 0.236 to 0.277.<\/jats:p>","DOI":"10.3389\/frai.2025.1516295","type":"journal-article","created":{"date-parts":[[2025,6,27]],"date-time":"2025-06-27T11:48:53Z","timestamp":1751024933000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Data stream-pairwise bottleneck transformer for engagement estimation from video conversation"],"prefix":"10.3389","volume":"8","author":[{"given":"Keita","family":"Suzuki","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nobukatsu","family":"Hojo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kazutoshi","family":"Shinoda","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Saki","family":"Mizuno","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ryo","family":"Masumura","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2025,6,27]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/WACV.2016.7477553","article-title":"\u201cOpenFace: an open source facial behavior analysis toolkit,\u201d","author":"Baltru\u0161aitis","year":"2016","journal-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision"},{"key":"B2","doi-asserted-by":"publisher","DOI":"10.1145\/2401836.2401846","article-title":"\u201cGaze and conversational engagement in multiparty video conversation: an annotation scheme and classification of high and low levels of engagement,\u201d","author":"Bednarik","year":"2012","journal-title":"Proceedings of the 4th Workshop on Eye Gaze in Intelligent Human Machine Interaction"},{"key":"B3","doi-asserted-by":"publisher","first-page":"350","DOI":"10.1145\/3136755.3136780","article-title":"\u201cThe NoXi database: multimodal recordings of mediated novice-expert interactions,\u201d","author":"Cafaro","year":"2017","journal-title":"Proceedings of the International Conference on Multimodal Interaction"},{"key":"B4","doi-asserted-by":"publisher","first-page":"3345","DOI":"10.1109\/TAFFC.2022.3178689","article-title":"Dyadic affect in parent-child multi-modal interaction: introducing the DAMI-P2C dataset and its preliminary analysis","volume":"14","author":"Chen","year":"2022","journal-title":"IEEE Trans. Affect. Comput"},{"key":"B5","doi-asserted-by":"publisher","first-page":"2426","DOI":"10.21437\/Interspeech.2021-329","article-title":"\u201cUnsupervised cross-lingual representation learning for speech recognition,\u201d","author":"Conneau","year":"2021","journal-title":"Proceedings of the Annual Conference of the International Speech Communication Association"},{"key":"B6","doi-asserted-by":"publisher","first-page":"440","DOI":"10.1145\/3340555.3353765","article-title":"\u201cEngagement modeling in dyadic interaction,\u201d","author":"Dermouche","year":"2019","journal-title":"Proceedings of the International Conference on Multimodal Interaction"},{"key":"B7","doi-asserted-by":"publisher","first-page":"546","DOI":"10.1145\/3340555.3355710","article-title":"\u201cEmotiW 2019: automatic emotion, engagement and cohesion prediction tasks,\u201d","author":"Dhall","year":"2019","journal-title":"Proceedings of the International Conference on Multimodal Interaction"},{"key":"B8","doi-asserted-by":"publisher","first-page":"653","DOI":"10.1145\/3242969.3264993","article-title":"\u201cEmotiW 2018: audio-video, student engagement and group-level affect prediction tasks,\u201d","author":"Dhall","year":"2018","journal-title":"Proceedings of the International Conference on Multimodal Interaction"},{"key":"B9","doi-asserted-by":"publisher","first-page":"59","DOI":"10.3102\/00346543074001059","article-title":"School engagement: potential of the concept, state of the evidence","volume":"74","author":"Fredricks","year":"2004","journal-title":"Rev. Educ. Res"},{"key":"B10","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.learninstruc.2016.02.002","article-title":"Student engagement, context, and adjustment: addressing definitional, measurement, and methodological issues","volume":"43","author":"Fredricks","year":"2016","journal-title":"Learn. Instr"},{"key":"B11","doi-asserted-by":"publisher","first-page":"42","DOI":"10.1145\/2663204.2663264","article-title":"\u201cThe additive value of multimodal features for predicting engagement, frustration, and learning during tutoring,\u201d","author":"Grafsgaard","year":"2014","journal-title":"Proceedings of the International Conference on Multimodal Interaction"},{"key":"B12","article-title":"DAiSEE: towards user engagement recognition in the wild","author":"Gupta","year":"2016","journal-title":"arXiv preprint arXiv:1609.01885"},{"key":"B13","doi-asserted-by":"publisher","first-page":"770","DOI":"10.1109\/CVPR.2016.90","article-title":"\u201cDeep residual learning for image recognition,\u201d","author":"He","year":"2016","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"B14","doi-asserted-by":"publisher","first-page":"567","DOI":"10.1145\/3340555.3355714","article-title":"\u201cEngagement intensity prediction with facial behavior features,\u201d","author":"Huynh","year":"2019","journal-title":"Proceedings of the International Conference on Multimodal Interaction"},{"key":"B15","doi-asserted-by":"publisher","first-page":"314","DOI":"10.1145\/3577190.3614122","article-title":"\u201cHIINT: historical, intra-and inter-personal dynamics modeling with cross-person memory transformer,\u201d","author":"Kim","year":"2023","journal-title":"Proceedings of the International Conference on Multimodal Interaction"},{"key":"B16","doi-asserted-by":"publisher","first-page":"11383","DOI":"10.1145\/3664647.3688986","article-title":"\u201cTowards engagement prediction: a cross-modality dual-pipeline approach using visual and audio features,\u201d","author":"Kumar","year":"2024","journal-title":"Proceedings of the International Conference on Multimedia"},{"key":"B17","doi-asserted-by":"publisher","first-page":"3893","DOI":"10.24963\/ijcai.2023\/433","article-title":"\u201cMultipar-T: multiparty-transformer for capturing contingent behaviors in group conversations,\u201d","author":"Lee","year":"2023","journal-title":"Proceedings of the International Joint Conference on Artificial Intelligence"},{"key":"B18","doi-asserted-by":"publisher","first-page":"18","DOI":"10.1007\/978-3-030-58555-6_2","article-title":"\u201cGraph-based social relation reasoning,\u201d","author":"Li","year":"2020","journal-title":"Proceedings of the European Conference on Computer Vision"},{"key":"B19","doi-asserted-by":"publisher","first-page":"3312","DOI":"10.1109\/ICIP.2019.8803488","article-title":"\u201cFeature fusion of face and body for engagement intensity detection,\u201d","author":"Li","year":"2019","journal-title":"Proceeding of the IEEE International Conference on Image Processing"},{"key":"B20","doi-asserted-by":"publisher","first-page":"2980","DOI":"10.1109\/ICCV.2017.324","article-title":"\u201cFocal loss for dense object detection,\u201d","author":"Lin","year":"2017","journal-title":"Proceedings of the IEEE International Conference on Computer Vision"},{"key":"B21","doi-asserted-by":"publisher","first-page":"8044","DOI":"10.1109\/ICASSP40776.2020.9053308","article-title":"\u201cPredicting performance outcome with a conversational graph convolutional network for small group interactions,\u201d","author":"Lin","year":"2020","journal-title":"Proceedings of the International Conference on Acoustics, Speech and Signal Processing"},{"key":"B22","article-title":"\u201cOn the variance of the adaptive learning rate and beyond,\u201d","author":"Liu","year":"2020","journal-title":"Proceedings of the International Conference on Learning Representations"},{"key":"B23","doi-asserted-by":"publisher","first-page":"2782","DOI":"10.24963\/ijcai.2021\/383","article-title":"\u201cHierarchical temporal multi-instance learning for video-based student learning engagement assessment,\u201d","author":"Ma","year":"2021","journal-title":"Proceedings of the International Joint Conference on Artificial Intelligence"},{"key":"B24","first-page":"14200","article-title":"\u201cAttention bottlenecks for multimodal fusion,\u201d","author":"Nagrani","year":"2021","journal-title":"Proceedings of the International Conference on Neural Information Processing Systems"},{"key":"B25","doi-asserted-by":"publisher","first-page":"746","DOI":"10.1109\/TG.2023.3348230","article-title":"Video-based engagement estimation of game streamers: an interpretable multimodal neural network approach","volume":"16","author":"Pan","year":"2023","journal-title":"IEEE Trans. Games"},{"key":"B26","article-title":"YOLOv3: an incremental improvement","author":"Redmon","year":"2018","journal-title":"arXiv preprint arXiv:1804.02767"},{"key":"B27","first-page":"2517","article-title":"\u201cRoomReader: a multimodal corpus of online multiparty conversational interactions,\u201d","author":"Reverdy","year":"2022","journal-title":"Proceedings of the International Conference on Language Resources and Evaluation Conference"},{"key":"B28","doi-asserted-by":"publisher","DOI":"10.1145\/1734454.1734580","article-title":"\u201cRecognizing engagement in human-robot interaction,\u201d","author":"Rich","year":"2010","journal-title":"Proceedings of the ACM\/IEEE International Conference on Human-Robot Interaction"},{"key":"B29","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/FG.2013.6553805","article-title":"\u201cIntroducing the RECOLA multimodal corpus of remote collaborative and affective interactions,\u201d","author":"Ringeval","year":"2013","journal-title":"Proceedings of the IEEE International Conference and Workshops on Automatic Face and Gesture Recognition"},{"key":"B30","doi-asserted-by":"publisher","first-page":"305","DOI":"10.1145\/1957656.1957781","article-title":"\u201cAutomatic analysis of affective postures and body motion to detect engagement with a game companion,\u201d","author":"Sanghvi","year":"2011","journal-title":"Proceedings of the ACM\/IEEE International Conference on Human-Robot Interaction"},{"key":"B31","doi-asserted-by":"publisher","first-page":"2132","DOI":"10.1109\/TAFFC.2022.3188390","article-title":"\u201cClassifying emotions and engagement in online learning based on a single facial expression recognition neural network","volume":"13","author":"Savchenko","year":"2022","journal-title":"IEEE Trans. Affect. Comput."},{"key":"B32","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1080\/00461520.2014.1002924","article-title":"The challenges of defining and measuring student engagement in science","volume":"50","author":"Sinatra","year":"2015","journal-title":"Educ. Psychol"},{"key":"B33","doi-asserted-by":"publisher","first-page":"174","DOI":"10.1145\/3577190.3614164","article-title":"\u201cDo I have your attention: a large scale engagement prediction dataset and baselines,\u201d","author":"Singh","year":"2023","journal-title":"Proceedings of the International Conference on Multimodal Interaction"},{"key":"B34","doi-asserted-by":"publisher","first-page":"25228","DOI":"10.1109\/ACCESS.2024.3353053","article-title":"Multimodal engagement recognition from image traits using deep learning techniques","volume":"12","author":"Sukumaran","year":"2024","journal-title":"IEEE Access"},{"key":"B35","doi-asserted-by":"publisher","first-page":"309","DOI":"10.1109\/TAFFC.2023.3274829","article-title":"Efficient multimodal transformer with dual-level feature restoration for robust multimodal sentiment analysis","volume":"15","author":"Sun","year":"2023","journal-title":"IEEE Trans. Affect. Comput"},{"key":"B36","doi-asserted-by":"publisher","first-page":"4079","DOI":"10.21437\/Interspeech.2024-1329","article-title":"\u201cParticipant-pair-wise bottleneck transformer for engagement estimation from video conversation,\u201d","author":"Suzuki","year":"2024","journal-title":"Proceedings of the Annual Conference of the International Speech Communication Association"},{"key":"B37","article-title":"\u201cFixing the train-test resolution discrepancy,\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Touvron","year":"2019"},{"key":"B38","first-page":"6000","article-title":"\u201cAttention is all you need,\u201d","author":"Vaswani","year":"2017","journal-title":"Proceedings of the International Conference on Neural Information Processing Systems"},{"key":"B39","doi-asserted-by":"publisher","first-page":"551","DOI":"10.1145\/3340555.3355711","article-title":"\u201cBootstrap model ensemble and rank loss for engagement intensity regression,\u201d","author":"Wang","year":"2019","journal-title":"Proceeding of the International Conference on Multimodal Interaction"},{"key":"B40","doi-asserted-by":"publisher","first-page":"9989","DOI":"10.1007\/s10639-023-12058-z","article-title":"CNN-Transformer: a deep learning method for automatically identifying learning engagement","volume":"15","author":"Xiong","year":"2023","journal-title":"Educ. Inf. Technol"},{"key":"B41","first-page":"1626","article-title":"\u201cGroup behavior recognition using attention- and graph-based neural networks,\u201d","author":"Yang","year":"2020","journal-title":"Proceedings of the European Conference on Artificial Intelligence"},{"key":"B42","doi-asserted-by":"publisher","first-page":"5525","DOI":"10.1109\/CVPR.2016.596","article-title":"\u201cWIDER FACE: a face detection benchmark,\u201d","author":"Yang","year":"2016","journal-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition"},{"key":"B43","doi-asserted-by":"publisher","first-page":"847","DOI":"10.1038\/s41562-019-0618-2","article-title":"A social interaction field model accurately identifies static and dynamic social groupings","volume":"3","author":"Zhou","year":"2019","journal-title":"Nat. Hum. Behav"}],"container-title":["Frontiers in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2025.1516295\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,27]],"date-time":"2025-06-27T11:48:56Z","timestamp":1751024936000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2025.1516295\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,27]]},"references-count":43,"alternative-id":["10.3389\/frai.2025.1516295"],"URL":"https:\/\/doi.org\/10.3389\/frai.2025.1516295","relation":{},"ISSN":["2624-8212"],"issn-type":[{"value":"2624-8212","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,27]]},"article-number":"1516295"}}