{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T16:06:16Z","timestamp":1781798776440,"version":"3.54.5"},"reference-count":45,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2019,10,26]],"date-time":"2019-10-26T00:00:00Z","timestamp":1572048000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["MTI"],"abstract":"<jats:p>We investigated the mouth-opening transition pattern (MOTP), which represents the change of mouth-opening degree during the end of an utterance, and used it to predict the next speaker and utterance interval between the start time of the next speaker\u2019s utterance and the end time of the current speaker\u2019s utterance in a multi-party conversation. We first collected verbal and nonverbal data that include speech and the degree of mouth opening (closed, narrow-open, wide-open) of participants that were manually annotated in four-person conversation. A key finding of the MOTP analysis is that the current speaker often keeps her mouth narrow-open during turn-keeping and starts to close it after opening it narrowly or continues to open it widely during turn-changing. The next speaker often starts to open her mouth narrowly after closing it during turn-changing. Moreover, when the current speaker starts to close her mouth after opening it narrowly in turn-keeping, the utterance interval tends to be short. In contrast, when the current speaker and the listeners open their mouths narrowly after opening them narrowly and then widely, the utterance interval tends to be long. On the basis of these results, we implemented prediction models of the next-speaker and utterance interval using MOTPs. As a multimodal-feature fusion, we also implemented models using eye-gaze behavior, which is one of the most useful items of information for prediction of next-speaker and utterance interval according to our previous study, in addition to MOTPs. The evaluation result of the models suggests that the MOTPs of the current speaker and listeners are effective for predicting the next speaker and utterance interval in multi-party conversation. Our multimodal-feature fusion model using MOTPs and eye-gaze behavior is more useful for predicting the next speaker and utterance interval than using only one or the other.<\/jats:p>","DOI":"10.3390\/mti3040070","type":"journal-article","created":{"date-parts":[[2019,10,28]],"date-time":"2019-10-28T04:44:31Z","timestamp":1572237871000},"page":"70","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":15,"title":["Prediction of Who Will Be Next Speaker and When Using Mouth-Opening Pattern in Multi-Party Conversation"],"prefix":"10.3390","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1025-1148","authenticated-orcid":false,"given":"Ryo","family":"Ishii","sequence":"first","affiliation":[{"name":"NTT Media Intelligence Laboratories, NTT Corporation, 1-1, Hikarinooka, Yokosuka-shi, Kanagawa 239-0847, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kazuhiro","family":"Otsuka","sequence":"additional","affiliation":[{"name":"NTT Communication Science Laboratories, NTT Corporation, 3-1, Morinosato Wakamiya, Atsugi-shi, Kanagawa 243-0198, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shiro","family":"Kumano","sequence":"additional","affiliation":[{"name":"NTT Communication Science Laboratories, NTT Corporation, 3-1, Morinosato Wakamiya, Atsugi-shi, Kanagawa 243-0198, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ryuichiro","family":"Higashinaka","sequence":"additional","affiliation":[{"name":"NTT Media Intelligence Laboratories, NTT Corporation, 1-1, Hikarinooka, Yokosuka-shi, Kanagawa 239-0847, Japan"},{"name":"NTT Communication Science Laboratories, NTT Corporation, 3-1, Morinosato Wakamiya, Atsugi-shi, Kanagawa 243-0198, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Junji","family":"Tomita","sequence":"additional","affiliation":[{"name":"NTT Media Intelligence Laboratories, NTT Corporation, 1-1, Hikarinooka, Yokosuka-shi, Kanagawa 239-0847, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,10,26]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Gatica-Perez, D. (2006, January 3\u20136). Analyzing group interactions in conversations: A review. Proceedings of the MFI, Heidelberg, Germany.","DOI":"10.1109\/MFI.2006.265658"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1109\/MSP.2011.941100","article-title":"Conversational scene analysis","volume":"28","author":"Otsuka","year":"2011","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Ishii, R., Kumano, S., and Otsuka, K. (2016, January 12\u201316). Multimodal Fusion using Respiration and Gaze for Predicting Next Speaker in Multi-Party Meetings. Proceedings of the ICMI, Tokyo, Japan.","DOI":"10.1145\/2993148.2993189"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Ishii, R., Kumano, S., and Otsuka, K. (2015, January 19\u201324). Predicting Next Speaker Using Head Movement in Multi-party Meetings. Proceedings of the ICASSP, Queensland, Australia.","DOI":"10.1109\/ICASSP.2015.7178385"},{"key":"ref_5","first-page":"4","article-title":"Predicting of Who Will Be the Next Speaker and When Using Gaze Behavior in Multiparty Meetings","volume":"6","author":"Ishii","year":"2016","journal-title":"ACM TiiS"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Ishii, R., Otsuka, K., Kumano, S., and Yamato, J. (2014, January 12\u201316). Analysis of Respiration for Prediction of Who Will Be Next Speaker and When?. Proceedings of the ICMI, Istanbul, Turkey.","DOI":"10.1145\/2663204.2663271"},{"key":"ref_7","first-page":"20","article-title":"Using Respiration to Predict Who Will Speak Next and When in Multiparty Meetings","volume":"6","author":"Ishii","year":"2016","journal-title":"ACM TiiS"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"6585","DOI":"10.1523\/JNEUROSCI.14-11-06585.1994","article-title":"Speech Motor Coordination and Control: Evidence from Lip, Jaw, and Laryngeal Movements","volume":"14","author":"Gracco","year":"1994","journal-title":"J. Neurosci."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"696","DOI":"10.1353\/lan.1974.0010","article-title":"A simplest systematics for the organisation of turn taking for conversation","volume":"50","author":"Sacks","year":"1974","journal-title":"Language"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1016\/0001-6918(67)90005-4","article-title":"Some functions of gaze direction in social interaction","volume":"26","author":"Kendon","year":"1967","journal-title":"Acta Psychol."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"495","DOI":"10.3389\/fpsyg.2015.00495","article-title":"Dutch and English toddlers\u2019 use of linguistic cues in predicting upcoming turn transitions","volume":"6","author":"Lammertink","year":"2015","journal-title":"Front. Psychol."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"6","DOI":"10.1016\/j.tics.2015.10.010","article-title":"Turn-taking in human communication\u2014Origins and implications for language processing","volume":"20","author":"Levinson","year":"2016","journal-title":"Trends Cogn. Sci."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Kawahara, T., Iwatate, T., and Takanashii, K. (2012, January 9\u201313). Prediction of turn-taking by combining prosodic and eye-gaze information in poster conversations. Proceedings of the INTERSPEECH, Portland, OR, USA.","DOI":"10.21437\/Interspeech.2012-226"},{"key":"ref_14","first-page":"12","article-title":"Gaze and turn-taking behavior in casual conversational interactions","volume":"3","author":"Jokinen","year":"2013","journal-title":"ACM TiiS"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Ishii, R., Otsuka, K., Kumano, S., Matsuda, M., and Yamato, J. (2013, January 9\u201313). Predicting Next Speaker and Timing from Gaze Transition Patterns in Multi-Party Meetings. Proceedings of the ICMI, Sydney, Australia.","DOI":"10.1145\/2522848.2522856"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Ishii, R., Otsuka, K., Kumano, S., and Yamato, J. (2014, January 4\u20139). Analysis and Modeling of Next Speaking Start Timing based on Gaze Behavior in Multi-party Meetings. Proceedings of the ICASSP, Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6853685"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"515","DOI":"10.3389\/fpsyg.2015.00098","article-title":"Unaddressed participants\u2019 gaze in multi-person interaction: optimizing recipiency","volume":"6","author":"Holler","year":"2015","journal-title":"Front. Psychol."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"54","DOI":"10.1080\/08351813.2017.1262143","article-title":"Eye blinking as addressee feedback in face-to-face conversation","volume":"50","author":"Holler","year":"2017","journal-title":"Res. Lang. Soc. Interact."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Ishii, R., Kumano, S., and Otsuka, K. (2017, January 17\u201320). Prediction of Next-Utterance Timing using Head Movement in Multi-Party Meetings. Proceedings of the HAI, Bielefeld, Germany.","DOI":"10.1145\/3125739.3125765"},{"key":"ref_20","first-page":"25","article-title":"Processing language in face-to-face conversation: Questons with gestures get faster responses","volume":"6","author":"Holler","year":"2018","journal-title":"Psychon. Bull. Rev."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Chen, L., and Harper, M.P. (2009, January 2\u20134). Multimodal floor control shift detection. Proceedings of the ICMI, Cambridge, MA, USA.","DOI":"10.1145\/1647314.1647320"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"de Kok, I., and Heylen, D. (2009, January 2\u20134). Multimodal end-of-turn prediction in multi-party meetings. Proceedings of the ICMI, Cambridge, MA, USA.","DOI":"10.1145\/1647314.1647332"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ferrer, L., Shriberg, E., and Stolcke, A. (2002, January 16\u201320). Is the speaker done yet? Faster and more accurate end-of-utterance detection using prosody in human-computer dialog. Proceedings of the INTERSPEECH, Denver, CO, USA.","DOI":"10.21437\/ICSLP.2002-565"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Laskowski, K., Edlund, J., and Heldner, M. (2011, January 22\u201327). A single-port non-parametric model of turn-taking in multi-party conversation. Proceedings of the ICASSP, Prague, Czech Republic.","DOI":"10.1109\/ICASSP.2011.5947629"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Schlangen, D. (2006, January 17\u201321). From reaction to prediction: experiments with computational models of turn-taking. Proceedings of the INTERSPEECH, Pittsburgh, PA, USA.","DOI":"10.21437\/Interspeech.2006-550"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Dielmann, A., Garau, G., and Bourlard, H. (2010, January 26\u201330). Floor holder detection and end of speaker turn prediction in meetings. Proceedings of the INTERSPEECH, Makuhari, Japan.","DOI":"10.21437\/Interspeech.2010-632"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Itoh, T., Kitaoka, N., and Nishimura, R. (2009, January 6\u201310). Subjective experiments on influence of response timing in spoken dialogues. Proceedings of the ISCA, Brighton, UK.","DOI":"10.21437\/Interspeech.2009-534"},{"key":"ref_28","unstructured":"Inoue, M., Yoroizawa, I., and Okubo, S. (1984). Human Factors Oriented Design Objectives for Video Teleconferencing Systems. ITS, 66\u201373."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"198","DOI":"10.1109\/34.982900","article-title":"Extraction of visual features for lipreading","volume":"24","author":"Matthews","year":"2002","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Chakravarty, P., Mirzaei, S., and Tuytelaars, T. (2015, January 9\u201313). Who\u2019s speaking?: Audio-supervised classification of active speakers in video. Proceedings of the ICMI, Seattle, WA, USA.","DOI":"10.1145\/2818346.2820780"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Chakravarty, P., Zegers, J., Tuytelaars, T., and hamme, H.V. (2016, January 12\u201316). Active speaker detection with audio-visual co-training. Proceedings of the ICMI, Tokyo, Japan.","DOI":"10.1145\/2993148.2993172"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Cech, J., Mittal, R., Deleforge, A., Sanchez-Riera, J., AlamedaPineda, X., and Horaud, R. (2013, January 15\u201317). Active-speaker detection and localization with microphones and cameras embedded into a robotic head. Proceedings of the Humanoids, Atlanta, GA, USA.","DOI":"10.1109\/HUMANOIDS.2013.7029977"},{"key":"ref_33","unstructured":"Cutler, R., and Davis, L. (August, January 30). Look who\u2019s talking: Speaker detection using video and audio correlation. Proceedings of the ICME, New York, NY, USA."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Haider, F., Luz, S., and Campbell, N. (2016, January 7\u20139). Active speaker detection in human machine multiparty dialogue using visual prosody information. Proceedings of the GlobalSIP, Washington, DC, USA.","DOI":"10.1109\/GlobalSIP.2016.7906033"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Haider, F., Luz, S., Vogel, C., and Campbell, N. (2018, January 2\u20136). Improving Response Time of Active Speaker Detection using Visual Prosody Information Prior to Articulation. Proceedings of the INTERSPEECH, Hyderabad, India.","DOI":"10.21437\/Interspeech.2018-2310"},{"key":"ref_36","unstructured":"Murai, K. (2011). Speaker Predicting Apparatus, Speaker Predicting Method, and Program Product for Predicting Speaker. (20070120966), U.S. Patent."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"3336","DOI":"10.1016\/j.patcog.2012.02.024","article-title":"A local region based approach to lip tracking","volume":"45","author":"Cheunga","year":"2012","journal-title":"Pattern Recognit."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"295","DOI":"10.1177\/002383099804100404","article-title":"An analysis of turn-taking and backchannels based on prosodic and syntactic features in Japanese Map Task dialogs","volume":"41","author":"Koiso","year":"1998","journal-title":"Lang. Speech"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Ekman, P., and Friesen, W.V. (1978). The Facial Action Coding System: A Technique for the Measurement of Facial Movement, Consulting Psychologists Press.","DOI":"10.1037\/t27734-000"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"322","DOI":"10.1037\/0033-2909.88.2.322","article-title":"Integration and generalisation of Kappas for multiple raters","volume":"88","author":"Conger","year":"1980","journal-title":"Psychol. Bull."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Otsuka, K., Araki, S., Mikami, D., Ishizuka, K., Fujimoto, M., and Yamato, J. (2009, January 2\u20134). Realtime meeting analysis and 3D meeting viewer based on omnidirectional multimodal sensors. Proceedings of the ICMI, Cambridge, MA, USA.","DOI":"10.1145\/1647314.1647354"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"205","DOI":"10.2307\/2529686","article-title":"The analysis of residuals in cross-classified tables","volume":"29","author":"Haberman","year":"1973","journal-title":"Biometrics"},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"637","DOI":"10.1162\/089976601300014493","article-title":"Improvements to Platt\u2019s SMO Algorithm for SVM Classifier Design","volume":"13","author":"Keerthi","year":"2001","journal-title":"Neural Comput."},{"key":"ref_44","first-page":"2533","article-title":"WEKA\u2013Experiences with a Java Open-Source Project","volume":"11","author":"Bouckaert","year":"2010","journal-title":"J. Mach. Learn. Res."},{"key":"ref_45","unstructured":"Amos, B., Ludwiczuk, B., and Satyanarayanan, M. (2016). OpenFace: A General-Purpose Face Recognition Library with Mobile Applications, CMU School of Computer Science. Technical Report, CMU-CS-16-118."}],"container-title":["Multimodal Technologies and Interaction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2414-4088\/3\/4\/70\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:29:38Z","timestamp":1760189378000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2414-4088\/3\/4\/70"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,10,26]]},"references-count":45,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2019,12]]}},"alternative-id":["mti3040070"],"URL":"https:\/\/doi.org\/10.3390\/mti3040070","relation":{},"ISSN":["2414-4088"],"issn-type":[{"value":"2414-4088","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,10,26]]}}}