{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T13:19:48Z","timestamp":1771679988743,"version":"3.50.1"},"reference-count":26,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2016,5,5]],"date-time":"2016-05-05T00:00:00Z","timestamp":1462406400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Interact. Intell. Syst."],"published-print":{"date-parts":[[2016,5,5]]},"abstract":"<jats:p>In multiparty meetings, participants need to predict the end of the speaker\u2019s utterance and who will start speaking next, as well as consider a strategy for good timing to speak next. Gaze behavior plays an important role in smooth turn-changing. This article proposes a prediction model that features three processing steps to predict (I) whether turn-changing or turn-keeping will occur, (II) who will be the next speaker in turn-changing, and (III) the timing of the start of the next speaker\u2019s utterance. For the feature values of the model, we focused on gaze transition patterns and the timing structure of eye contact between a speaker and a listener near the end of the speaker\u2019s utterance. Gaze transition patterns provide information about the order in which gaze behavior changes. The timing structure of eye contact is defined as who looks at whom and who looks away first, the speaker or listener, when eye contact between the speaker and a listener occurs. We collected corpus data of multiparty meetings, using the data to demonstrate relationships between gaze transition patterns and timing structure and situations (I), (II), and (III). The results of our analyses indicate that the gaze transition pattern of the speaker and listener and the timing structure of eye contact have a strong association with turn-changing, the next speaker in turn-changing, and the start time of the next utterance. On the basis of the results, we constructed prediction models using the gaze transition patterns and timing structure. The gaze transition patterns were found to be useful in predicting turn-changing, the next speaker in turn-changing, and the start time of the next utterance. Contrary to expectations, we did not find that the timing structure is useful for predicting the next speaker and the start time. This study opens up new possibilities for predicting the next speaker and the timing of the next utterance using gaze transition patterns in multiparty meetings.<\/jats:p>","DOI":"10.1145\/2757284","type":"journal-article","created":{"date-parts":[[2016,5,5]],"date-time":"2016-05-05T13:23:22Z","timestamp":1462454602000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":36,"title":["Prediction of Who Will Be the Next Speaker and When Using Gaze Behavior in Multiparty Meetings"],"prefix":"10.1145","volume":"6","author":[{"given":"Ryo","family":"Ishii","sequence":"first","affiliation":[{"name":"NTT Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kazuhiro","family":"Otsuka","sequence":"additional","affiliation":[{"name":"NTT Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shiro","family":"Kumano","sequence":"additional","affiliation":[{"name":"NTT Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junji","family":"Yamato","sequence":"additional","affiliation":[{"name":"NTT Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,5,5]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/1756006.1953016"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1647314.1647320"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1037\/0033-2909.88.2.322"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1647314.1647332"},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the Annual Conference on the International Speech Communication Association. 2306--2309","author":"Dielmann Alfred","year":"2010","unstructured":"Alfred Dielmann , Giulia Garau , and Herv\u00e9 Bourlard . 2010 . Floor holder detection and end of speaker turn prediction in meetings . In Proceedings of the Annual Conference on the International Speech Communication Association. 2306--2309 . Alfred Dielmann, Giulia Garau, and Herv\u00e9 Bourlard. 2010. Floor holder detection and end of speaker turn prediction in meetings. In Proceedings of the Annual Conference on the International Speech Communication Association. 2306--2309."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1037\/h0033031"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the Annual Conference on the International Speech Communication Association","volume":"3","author":"Ferrer Luciana","year":"2002","unstructured":"Luciana Ferrer , Elizabeth Shriberg , and Andreas Stolcke . 2002 . Is the speaker done yet? Faster and more accurate end-of-utterance detection using prosody in human--computer dialog . In Proceedings of the Annual Conference on the International Speech Communication Association , Vol. 3 . 2061--2064. Luciana Ferrer, Elizabeth Shriberg, and Andreas Stolcke. 2002. Is the speaker done yet? Faster and more accurate end-of-utterance detection using prosody in human--computer dialog. In Proceedings of the Annual Conference on the International Speech Communication Association, Vol. 3. 2061--2064."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/MFI.2006.265658"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.2307\/2529686"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the International Conference on Autonomous Agents and Multi-Agent Systems.","author":"Huang Lixing","year":"2011","unstructured":"Lixing Huang , Louis-Philippe Morency , and Jonathan Gratch . 2011 . A multimodal end-of-turn prediction model: Learning from para social consensus sampling . In Proceedings of the International Conference on Autonomous Agents and Multi-Agent Systems. Lixing Huang, Louis-Philippe Morency, and Jonathan Gratch. 2011. A multimodal end-of-turn prediction model: Learning from para social consensus sampling. In Proceedings of the International Conference on Autonomous Agents and Multi-Agent Systems."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2499474.2499481"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the Conference of the European Chapter of the Association for Computational Linguistics.","author":"Jovanovic Natasa","unstructured":"Natasa Jovanovic , Rieks op den Akker, and Anton Nijholt. 2006. Addressee identification in face-to-face meetings . In Proceedings of the Conference of the European Chapter of the Association for Computational Linguistics. Natasa Jovanovic, Rieks op den Akker, and Anton Nijholt. 2006. Addressee identification in face-to-face meetings. In Proceedings of the Conference of the European Chapter of the Association for Computational Linguistics."},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the Annual Conference on the International Speech Communication Association.","author":"Kawahara Tatsuya","year":"2012","unstructured":"Tatsuya Kawahara , Takuma Iwatate , and Katsuya Takanashii . 2012 . Prediction of turn-taking by combining prosodic and eye-gaze information in poster conversations . In Proceedings of the Annual Conference on the International Speech Communication Association. Tatsuya Kawahara, Takuma Iwatate, and Katsuya Takanashii. 2012. Prediction of turn-taking by combining prosodic and eye-gaze information in poster conversations. In Proceedings of the Annual Conference on the International Speech Communication Association."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1162\/089976601300014493"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1016\/0001-6918(67)90005-4"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1177\/002383099804100404"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2011.5947629"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the SIGHAN Workshop on Chinese Language Processing.","author":"Levow Gina-Anne","year":"2005","unstructured":"Gina-Anne Levow . 2005 . Turn-taking in Mandarin dialogue: Interactions of tones and intonation . In Proceedings of the SIGHAN Workshop on Chinese Language Processing. Gina-Anne Levow. 2005. Turn-taking in Mandarin dialogue: Interactions of tones and intonation. In Proceedings of the SIGHAN Workshop on Chinese Language Processing."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2014.02.002"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-85483-8_18"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2011.941100"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1647314.1647354"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1353\/lan.1974.0010"},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of the Annual Conference on the International Speech Communication Association. 17--21","author":"Schlangen David","year":"2006","unstructured":"David Schlangen . 2006 . From reaction to prediction experiments with computational models of turn-taking . In Proceedings of the Annual Conference on the International Speech Communication Association. 17--21 . David Schlangen. 2006. From reaction to prediction experiments with computational models of turn-taking. In Proceedings of the Annual Conference on the International Speech Communication Association. 17--21."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:STCO.0000035301.49549.88"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4757-3264-1"}],"container-title":["ACM Transactions on Interactive Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2757284","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2757284","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T06:16:27Z","timestamp":1750227387000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2757284"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,5,5]]},"references-count":26,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2016,5,5]]}},"alternative-id":["10.1145\/2757284"],"URL":"https:\/\/doi.org\/10.1145\/2757284","relation":{},"ISSN":["2160-6455","2160-6463"],"issn-type":[{"value":"2160-6455","type":"print"},{"value":"2160-6463","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,5,5]]},"assertion":[{"value":"2014-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-05-05","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}