{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T13:49:53Z","timestamp":1767275393297,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":50,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,9,14]],"date-time":"2021-09-14T00:00:00Z","timestamp":1631577600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2021,9,14]]},"DOI":"10.1145\/3472306.3478360","type":"proceedings-article","created":{"date-parts":[[2021,9,10]],"date-time":"2021-09-10T10:15:37Z","timestamp":1631268937000},"page":"131-138","source":"Crossref","is-referenced-by-count":15,"title":["Multimodal and Multitask Approach to Listener's Backchannel Prediction"],"prefix":"10.1145","author":[{"given":"Ryo","family":"Ishii","sequence":"first","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xutong","family":"Ren","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michal","family":"Muszynski","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Louis-Philippe","family":"Morency","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,9,14]]},"reference":[{"volume-title":"Yao Chong Lim, and Louis-Philippe Morency","year":"2018","author":"Baltrusaitis Tadas","key":"e_1_3_2_1_1_1"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"crossref","unstructured":"P. Blache Massina Abderrahmane S. Rauzy and R. Bertrand. 2020. An integrated model for predicting backchannel feedbacks. In IVA.  P. Blache Massina Abderrahmane S. Rauzy and R. Bertrand. 2020. An integrated model for predicting backchannel feedbacks. In IVA.","DOI":"10.1145\/3383652.3423948"},{"volume-title":"Harper","year":"2009","author":"Chen Lei","key":"e_1_3_2_1_3_1"},{"key":"e_1_3_2_1_4_1","unstructured":"Kyunghyun Cho Bart van Merrienboer \u00c7aglar G\u00fcl\u00e7ehre Dzmitry Bahdanau Fethi Bougares Holger Schwenk and Yoshua Bengio. 2014. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. In EMNLP. 1724--1734.  Kyunghyun Cho Bart van Merrienboer \u00c7aglar G\u00fcl\u00e7ehre Dzmitry Bahdanau Fethi Bougares Holger Schwenk and Yoshua Bengio. 2014. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. In EMNLP. 1724--1734."},{"volume-title":"Turn-taking in Human Communication - Origins and Implications for Language Processing. Trends in cognitive sciences 20","year":"2016","author":"Levinson Stephen C.","key":"e_1_3_2_1_5_1"},{"volume-title":"Multimodal End-of-turn Prediction in Multi-party Meetings. In ICMI. 91--98","year":"2009","author":"de Kok Iwan","key":"e_1_3_2_1_6_1"},{"volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL. 4171--4186.","year":"2019","author":"Devlin Jacob","key":"e_1_3_2_1_7_1"},{"volume-title":"Floor Holder Detection and End of Speaker Turn Prediction in Meetings. In INTERSPEECH. 2306--2309","year":"2010","author":"Dielmann Alfred","key":"e_1_3_2_1_8_1"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"crossref","unstructured":"Florian Eyben Felix Weninger Florian Gross and Bj\u00f6rn Schuller. 2013. Recent Developments in OpenSMILE the Munich Open-Source Multimedia Feature Extractor. In ACM MM. 835--838.  Florian Eyben Felix Weninger Florian Gross and Bj\u00f6rn Schuller. 2013. Recent Developments in OpenSMILE the Munich Open-Source Multimedia Feature Extractor. In ACM MM. 835--838.","DOI":"10.1145\/2502081.2502224"},{"key":"e_1_3_2_1_10_1","first-page":"2061","article-title":"Is the Speaker Done Yet? Faster and More Accurate End-of-utterance Detection using Prosody in Human-computer Dialog","volume":"3","author":"Ferrer Luciana","year":"2002","journal-title":"INTERSPEECH"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Shinya Fujie Kenta Fukushima and Tetsunori Kobayashi. 2005. Back-channel feedback generation using linguistic and nonlinguistic information and its application to spoken dialogue system. In INTERSPEECH. 889--892.  Shinya Fujie Kenta Fukushima and Tetsunori Kobayashi. 2005. Back-channel feedback generation using linguistic and nonlinguistic information and its application to spoken dialogue system. In INTERSPEECH. 889--892.","DOI":"10.21437\/Interspeech.2005-400"},{"volume-title":"Audio Set: An Ontology and Human-labeled Dataset for Audio Events. In ICASSP. 776--780.","year":"2017","author":"Gemmeke Jort F.","key":"e_1_3_2_1_12_1"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Kohei Hara Koji Inoue Katsuya Takanashi and Tatsuya Kawahara. 2018. Prediction of Turn-taking Using Multitask Learning with Prediction of Backchannels and Fillers. In INTERSPEECH. 991--995.  Kohei Hara Koji Inoue Katsuya Takanashi and Tatsuya Kawahara. 2018. Prediction of Turn-taking Using Multitask Learning with Prediction of Backchannels and Fillers. In INTERSPEECH. 991--995.","DOI":"10.21437\/Interspeech.2018-1442"},{"key":"e_1_3_2_1_14_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In CVPR. 770--778.  Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In CVPR. 770--778."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"crossref","unstructured":"Shawn Hershey Sourish Chaudhuri Daniel P. W. Ellis Jort F. Gemmeke Aren Jansen Channing Moore Manoj Plakal Devin Platt Rif A. Saurous Bryan Seybold Malcolm Slaney Ron Weiss and Kevin Wilson. 2017. CNN Architectures for Large-Scale Audio Classification. In ICASSP. 131--135.  Shawn Hershey Sourish Chaudhuri Daniel P. W. Ellis Jort F. Gemmeke Aren Jansen Channing Moore Manoj Plakal Devin Platt Rif A. Saurous Bryan Seybold Malcolm Slaney Ron Weiss and Kevin Wilson. 2017. CNN Architectures for Large-Scale Audio Classification. In ICASSP. 131--135.","DOI":"10.1109\/ICASSP.2017.7952132"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.3389\/fpsyg.2015.00098"},{"key":"e_1_3_2_1_17_1","first-page":"25","article-title":"Processing language in face-to-face conversation: Questons with gestures get faster responses","volume":"6","author":"Holler Judith","year":"2018","journal-title":"Psychonomic Bulletin Review"},{"key":"e_1_3_2_1_18_1","first-page":"1265","article-title":"Parasocial consensus sampling: Combining multiple perspectives to learn virtual human behavior","volume":"2","author":"Huang Lixing","year":"2010","journal-title":"AAMAS"},{"key":"e_1_3_2_1_19_1","unstructured":"Lixing Huang Louis-Philippe Morency and Jonathan Gratch. 2011. A Multimodal End-of-Turn Prediction Model: Learning from Parasocial Consensus Sampling. In AAMAS.  Lixing Huang Louis-Philippe Morency and Jonathan Gratch. 2011. A Multimodal End-of-Turn Prediction Model: Learning from Parasocial Consensus Sampling. In AAMAS."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1080\/08351813.2017.1262143"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Ryo Ishii Shiro Kumano and Kazuhiro Otsuka. 2015. Multimodal Fusion using Respiration and Gaze for Predicting Next Speaker in Multi-Party Meetings. In ICMI. 99--106.  Ryo Ishii Shiro Kumano and Kazuhiro Otsuka. 2015. Multimodal Fusion using Respiration and Gaze for Predicting Next Speaker in Multi-Party Meetings. In ICMI. 99--106.","DOI":"10.1145\/2818346.2820755"},{"volume-title":"Predicting Next Speaker Using Head Movement in Multi-party Meetings. In ICASSP. 2319--2323","year":"2015","author":"Ishii Ryo","key":"e_1_3_2_1_22_1"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"crossref","unstructured":"Ryo Ishii Shiro Kumano and Kazuhiro Otsuka. 2017. Prediction of Next-Utterance Timing using Head Movement in Multi-Party Meetings. In HAI. 181--187.  Ryo Ishii Shiro Kumano and Kazuhiro Otsuka. 2017. Prediction of Next-Utterance Timing using Head Movement in Multi-Party Meetings. In HAI. 181--187.","DOI":"10.1145\/3125739.3125765"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.3390\/mti3040070"},{"key":"e_1_3_2_1_25_1","first-page":"1","volume-title":"Predicting of Who Will Be the Next Speaker and When Using Gaze Behavior in Multiparty Meetings. ACM TiiS 6","author":"Ishii Ryo","year":"2016"},{"key":"e_1_3_2_1_26_1","first-page":"20","article-title":"Using Respiration to Predict Who Will Speak Next and When in Multiparty Meetings","volume":"6","author":"Ishii Ryo","year":"2016","journal-title":"ACM TiiS"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"crossref","unstructured":"Ryo Ishii Xutong Ren Michal Muszynski and Louis-Philippe Morency. 2020. Can Prediction of Turn-Management Willingness Improve Turn-Changing Modeling?. In IVA.  Ryo Ishii Xutong Ren Michal Muszynski and Louis-Philippe Morency. 2020. Can Prediction of Turn-Management Willingness Improve Turn-Changing Modeling?. In IVA.","DOI":"10.1145\/3383652.3423907"},{"key":"e_1_3_2_1_28_1","first-page":"12","article-title":"Gaze and turn-taking behavior in casual conversational interactions","volume":"3","author":"Jokinen Kristiina","year":"2013","journal-title":"ACM TiiS"},{"volume-title":"Measuring Emotional Expression with the Linguistic Inquiry and Word Count. J. psychology 120","year":"2007","author":"Kahn Jeffrey","key":"e_1_3_2_1_29_1"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"crossref","unstructured":"Tatsuya Kawahara Takuma Iwatate and Katsuya Takanashii. 2012. Prediction of Turn-taking by Combining Prosodic and Eye-gaze Information in Poster Conversations. In INTERSPEECH. 726--729.  Tatsuya Kawahara Takuma Iwatate and Katsuya Takanashii. 2012. Prediction of Turn-taking by Combining Prosodic and Eye-gaze Information in Poster Conversations. In INTERSPEECH. 726--729.","DOI":"10.21437\/Interspeech.2012-226"},{"volume-title":"Kingma and Jimmy Ba","year":"2015","author":"Diederik","key":"e_1_3_2_1_31_1"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1527\/tjsai.20.220"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1177\/002383099804100404"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"crossref","unstructured":"Divesh Lala Koji Inoue and Tatsuya Kawahara. 2018. Evaluation of Real-Time Deep Learning Turn-Taking Models for Multiple Dialogue Scenarios. In ICMI. 78--86.  Divesh Lala Koji Inoue and Tatsuya Kawahara. 2018. Evaluation of Real-Time Deep Learning Turn-Taking Models for Multiple Dialogue Scenarios. In ICMI. 78--86.","DOI":"10.1145\/3242969.3242994"},{"volume-title":"Dutch and English Toddlers' Use of Linguistic Cues in Predicting Upcoming Turn Transitions. Frontiers in Psychology","year":"2015","author":"Lammertink Imme","key":"e_1_3_2_1_35_1"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"crossref","unstructured":"Kornel Laskowski Jens Edlund and Mattias Heldner. 2011. A single-port nonparametric model of turn-taking in multi-party conversation. In ICASSP. 5600--5603.  Kornel Laskowski Jens Edlund and Mattias Heldner. 2011. A single-port nonparametric model of turn-taking in multi-party conversation. In ICASSP. 5600--5603.","DOI":"10.1109\/ICASSP.2011.5947629"},{"volume-title":"Improving Speech-Based End-of-Turn Detection Via Cross-Modal Representation Learning with Punctuated Text Data. ASRU","year":"2019","author":"Masumura Ryo","key":"e_1_3_2_1_37_1"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"crossref","unstructured":"Ryo Masumura Tomohiro Tanaka Atsushi Ando Ryo Ishii Ryuichiro Higashinaka and Yushi Aono. 2018. Neural Dialogue Context Online End-of-Turn Detection. In SIGdial. 224--228.  Ryo Masumura Tomohiro Tanaka Atsushi Ando Ryo Ishii Ryuichiro Higashinaka and Yushi Aono. 2018. Neural Dialogue Context Online End-of-Turn Detection. In SIGdial. 224--228.","DOI":"10.18653\/v1\/W18-5024"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"crossref","unstructured":"Louis-Philippe Morency Iwan de Kok and Jonathan Gratch. 2008. Predicting Listener Backchannels: A Probabilistic Multimodal Approach. In IVA. 176--190.  Louis-Philippe Morency Iwan de Kok and Jonathan Gratch. 2008. Predicting Listener Backchannels: A Probabilistic Multimodal Approach. In IVA. 176--190.","DOI":"10.1007\/978-3-540-85483-8_18"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"crossref","unstructured":"Markus Mueller David Leuschner Lars Briem Maria Schmidt Kevin Kilgour Sebastian Stueker and Alex Waibel. 2015. Using Neural Networks for Data-Driven Backchannel Prediction: A Survey on Input Features and Training Techniques. In Human-Computer Interaction: Interaction Technologies. 329--340.  Markus Mueller David Leuschner Lars Briem Maria Schmidt Kevin Kilgour Sebastian Stueker and Alex Waibel. 2015. Using Neural Networks for Data-Driven Backchannel Prediction: A Survey on Input Features and Training Techniques. In Human-Computer Interaction: Interaction Technologies. 329--340.","DOI":"10.1007\/978-3-319-20916-6_31"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"crossref","unstructured":"Matthew Roddy Gabriel Skantze and Naomi Harte. 2018. Multimodal Continuous Turn-Taking Prediction Using Multiscale RNNs. In ICMI. 186--190.  Matthew Roddy Gabriel Skantze and Naomi Harte. 2018. Multimodal Continuous Turn-Taking Prediction Using Multiscale RNNs. In ICMI. 186--190.","DOI":"10.1145\/3242969.3242997"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"crossref","unstructured":"Robin Ruede Markus M\u00fcller Sebastian St\u00fcker and Alex Waibel. 2019. Yeah Right Uh-Huh: A Deep Learning Backchannel Predictor. 247--258.  Robin Ruede Markus M\u00fcller Sebastian St\u00fcker and Alex Waibel. 2019. Yeah Right Uh-Huh: A Deep Learning Backchannel Predictor. 247--258.","DOI":"10.1007\/978-3-319-92108-2_25"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"crossref","unstructured":"David Schlangen. 2006. From Reaction to Prediction: Experiments with Computational Models of Turn-taking. In INTERSPEECH. 17--21.  David Schlangen. 2006. From Reaction to Prediction: Experiments with Computational Models of Turn-taking. In INTERSPEECH. 17--21.","DOI":"10.21437\/Interspeech.2006-550"},{"key":"e_1_3_2_1_45_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In ICLR.  Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In ICLR."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"crossref","unstructured":"Mohammad Soleymani Kalin Stefanov Sin-Hwa Kang Jan Ondras and Jonathan Gratch. 2019. Multimodal Analysis and Estimation of Intimate Self-Disclosure. In ICMI. 59--68.  Mohammad Soleymani Kalin Stefanov Sin-Hwa Kang Jan Ondras and Jonathan Gratch. 2019. Multimodal Analysis and Estimation of Intimate Self-Disclosure. In ICMI. 59--68.","DOI":"10.1145\/3340555.3353737"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"crossref","unstructured":"Khiet P. Truong Ronald Poppe and Dirk Heylen. 2010. A rule-based backchannel prediction model using pitch and pause information.. In INTERSPEECH. ISCA.  Khiet P. Truong Ronald Poppe and Dirk Heylen. 2010. A rule-based backchannel prediction model using pitch and pause information.. In INTERSPEECH. ISCA.","DOI":"10.21437\/Interspeech.2010-59"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSLP.1996.607961"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"crossref","unstructured":"Nigel Ward Diego Aguirre Gerardo Cervantes and Olac Fuentes. 2018. Turn-Taking Predictions across Languages and Genres Using an LSTM Recurrent Neural Network. In SLT. 831--837.  Nigel Ward Diego Aguirre Gerardo Cervantes and Olac Fuentes. 2018. Turn-Taking Predictions across Languages and Genres Using an LSTM Recurrent Neural Network. In SLT. 831--837.","DOI":"10.1109\/SLT.2018.8639673"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0378-2166(99)00109-5"}],"event":{"name":"IVA '21: ACM International Conference on Intelligent Virtual Agents","sponsor":["SIGAI ACM Special Interest Group on Artificial Intelligence"],"location":"Virtual Event Japan","acronym":"IVA '21"},"container-title":["Proceedings of the 21th ACM International Conference on Intelligent Virtual Agents"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3472306.3478360","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3472306.3478360","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:17:36Z","timestamp":1750191456000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3472306.3478360"}},"subtitle":["Can Prediction of Turn-changing and Turn-management Willingness Improve Backchannel Modeling?"],"short-title":[],"issued":{"date-parts":[[2021,9,14]]},"references-count":50,"alternative-id":["10.1145\/3472306.3478360","10.1145\/3472306"],"URL":"https:\/\/doi.org\/10.1145\/3472306.3478360","relation":{},"subject":[],"published":{"date-parts":[[2021,9,14]]}}}