{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,18]],"date-time":"2026-01-18T08:48:23Z","timestamp":1768726103986,"version":"3.49.0"},"reference-count":50,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2020,4,29]],"date-time":"2020-04-29T00:00:00Z","timestamp":1588118400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["MTI"],"abstract":"<jats:p>In recent years, companies have been seeking communication skills from their employees. Increasingly more companies have adopted group discussions during their recruitment process to evaluate the applicants\u2019 communication skills. However, the opportunity to improve communication skills in group discussions is limited because of the lack of partners. To solve this issue as a long-term goal, the aim of this study is to build an autonomous robot that can participate in group discussions, so that its users can repeatedly practice with it. This robot, therefore, has to perform humanlike behaviors with which the users can interact. In this study, the focus was on the generation of two of these behaviors regarding the head of the robot. One is directing its attention to either of the following targets: the other participants or the materials placed on the table. The second is to determine the timings of the robot\u2019s nods. These generation models are considered in three situations: when the robot is speaking, when the robot is listening, and when no participant including the robot is speaking. The research question is: whether these behaviors can be generated end-to-end from and only from the features of peer participants. This work is based on a data corpus containing 2.5 h of the discussion sessions of 10 four-person groups. Multimodal features, including the attention of other participants, voice prosody, head movements, and speech turns extracted from the corpus, were used to train support vector machine models for the generation of the two behaviors. The performances of the generation models of attentional focus were in an F-measure range between 0.4 and 0.6. The nodding model had an accuracy of approximately 0.65. Both experiments were conducted in the setting of leave-one-subject-out cross validation. To measure the perceived naturalness of the generated behaviors, a subject experiment was conducted. In the experiment, the proposed models were compared. They were based on a data-driven method with two baselines: (1) a simple statistical model based on behavior frequency and (2) raw experimental data. The evaluation was based on the observation of video clips, in which one of the subjects was replaced by a robot performing head movements in the above-mentioned three conditions. The experimental results showed that there was no significant difference from original human behaviors in the data corpus and proved the effectiveness of the proposed models.<\/jats:p>","DOI":"10.3390\/mti4020015","type":"journal-article","created":{"date-parts":[[2020,4,29]],"date-time":"2020-04-29T13:23:45Z","timestamp":1588166625000},"page":"15","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Generation of Head Movements of a Robot Using Multimodal Features of Peer Participants in Group Discussion Conversation"],"prefix":"10.3390","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0376-3535","authenticated-orcid":false,"given":"Hung-Hsuan","family":"Huang","sequence":"first","affiliation":[{"name":"Faculty of Informatics, The University of Fukuchiyama, Fukuchiyama, Kyoto 620-0886, Japan"},{"name":"Center for Advanced Intelligence Project, RIKEN, Kyoto, Kyoto 606-8501, Japan"},{"name":"Graduate School of Informatics, Kyoto University, Kyoto, Kyoto 606-8501, Japan"},{"name":"College of Information Science and Engineering, Ritsumeikan University, Kusatsu, Shiga 525-8577, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Seiya","family":"Kimura","sequence":"additional","affiliation":[{"name":"Center for Advanced Intelligence Project, RIKEN, Kyoto, Kyoto 606-8501, Japan"},{"name":"College of Information Science and Engineering, Ritsumeikan University, Kusatsu, Shiga 525-8577, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kazuhiro","family":"Kuwabara","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Ritsumeikan University, Kusatsu, Shiga 525-8577, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Toyoaki","family":"Nishida","sequence":"additional","affiliation":[{"name":"Faculty of Informatics, The University of Fukuchiyama, Fukuchiyama, Kyoto 620-0886, Japan"},{"name":"Center for Advanced Intelligence Project, RIKEN, Kyoto, Kyoto 606-8501, Japan"},{"name":"Graduate School of Informatics, Kyoto University, Kyoto, Kyoto 606-8501, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,4,29]]},"reference":[{"key":"ref_1","unstructured":"Chollet, M., Ochs, M., and Pelachaud, C. (2014, January 26\u201331). Mining a Multimodal Corpus for Non-Verbal Signals Sequences Conveying Attitudes. Proceedings of the 9th Edition of the Language Resources and Evaluation Conference (LREC 2014), Reykjavik, Iceland."},{"key":"ref_2","unstructured":"Jones, H., Chollet, M., Ochs, M., Sabouret, N., and Pelachaud, C. (2014, January 5\u20139). Expressing social attitudes in virtual agents for social coaching. Proceedings of the Autonomous Agents and Multi-Agent Systems (AAMAS\u201914), Paris, France."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Baur, T., Damian, I., Gebhard, P., Porayska-Pomsta, K., and Andre, E. (2013, January 8\u201314). A Job Interview Simulation: Social Cue-based Interaction with a Virtual Character. Proceedings of the 2013 International Conference on Social Computing (SocialCom 2013), Washington, DC, USA.","DOI":"10.1109\/SocialCom.2013.39"},{"key":"ref_4","unstructured":"Traum, D. (2003, January 11\u201313). Issues in Multiparty Dialogues. Proceedings of the Advances in Agent Communication, International Workshop on Agent Communication Languages (ACL\u201903), Halifax, NS, Canada."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Huang, H.H., Baba, N., and Nakano, Y. (2011, January 14\u201318). Making Virtual Conversational Agent Aware of the Addressee of Users\u2019 Utterances in Multi-user Conversation from Nonverbal Information. Proceedings of the 13th International Conference on Multimodal Interaction (ICMI\u201911), Alicante, Spain.","DOI":"10.1145\/2070481.2070557"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Baba, N., Huang, H.H., and Nakano, Y. (2012, January 26). Addressee Identification for Human-Human-Agent Multiparty Conversations in Different Proxemics. Proceedings of the 4th Workshop on Eye Gaze in Intelligent Human Machine Interaction: Eye Gaze and Multimodality, Santa Monica, CA, USA.","DOI":"10.1145\/2401836.2401842"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Nakano, Y., Baba, N., Huang, H.H., and Hayashi, Y. (2013, January 9\u201313). Implementation and Evaluation of Multimodal Addressee Identification Mechanism for Multiparty Conversation Systems. Proceedings of the 15th International Conference on Multimodal Interaction (ICMI 2013), Sydney, Australia.","DOI":"10.1145\/2522848.2522872"},{"key":"ref_8","unstructured":"Jokinen, K., and Parkson, S. (2011, January 27\u201328). Synchrony and copying in conversational interactions. Proceedings of the 3rd Nordic Symposium on Multimodal Interaction, Helsinki, Finland."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"159","DOI":"10.1111\/1467-9280.00232","article-title":"Reflexive Joing Attention depends on Lateralized Cortical Connections","volume":"11","author":"Kingstone","year":"2000","journal-title":"Psychol. Sci."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"507","DOI":"10.3758\/BF03196306","article-title":"Are eyes special? It depends on how you look at it","volume":"9","author":"Ristic","year":"2002","journal-title":"Psychon. Bull. Rev."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"71","DOI":"10.1080\/13506280344000220","article-title":"Why does the gaze of others direct visual attention?","volume":"11","author":"Downing","year":"2004","journal-title":"Vis. Cogn."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Otsuka, K. (2011). Multimodal Conversation Scene Analysis for Understanding People\u2019s Communicative Behaviors in Face-to-Face Meetings. Symposium on Human Interface, Springer.","DOI":"10.1007\/978-3-642-21669-5_21"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1016\/0001-6918(67)90005-4","article-title":"Some functions of gaze direction in social interaction","volume":"26","author":"Kendon","year":"1967","journal-title":"Acta Psychol."},{"key":"ref_14","unstructured":"Argyle, M., and Cook, M. (1976). Gaze and Mutual Gaze, Cambridge University Press."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1037\/h0033031","article-title":"Some Signals and Rules for Taking Speaking Turns in Conversations","volume":"23","author":"Duncan","year":"1972","journal-title":"J. Personal. Psychol."},{"key":"ref_16","unstructured":"Vertegaal, R., Slagter, R., van der Veer, G., and Nijholt, A. (April, January 31). Eye gaze patterns in conversations: There is more to conversational agents than meets the eyes. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, Seattle, WA, USA."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Takemae, Y., Otsuka, K., and Mukawa, N. (2003, January 2\u20138). Video Cut Editing Rule Based on Participants\u2019 Gaze in Multiparty Conversation. Proceedings of the 11th ACM International Conference on Multimedia, Berkeley, CA, USA.","DOI":"10.1145\/957013.957077"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Katzenmaier, M., Stiefelhagen, R., and Schultz, T. (2004, January 13\u201315). Identifying the Addressee in Human-Human-Robot Interactions based on Head Pose and Speech. Proceedings of the 6th international conference on Multimodal interfaces (ICM 2004), State College, PA, USA.","DOI":"10.1145\/1027933.1027959"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Gratch, J., Okhmatovskaia, A., Lamothe, F., Marsella, S., Morales, M., van der Werf, R., and Morency, L.P. (2006, January 21\u201323). Virtual Rapport. Proceedings of the 6th International Conference on Intelligent Virtual Agents (IVA 2006), Marina Del Rey, CA, USA.","DOI":"10.1007\/11821830"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Ishii, C.T., Liu, C., Ishiguro, H., and Hagita, N. (2010, January 2\u20135). Head motions during dialogue speech and nod timing control in humanoid robots. Proceedings of the 5th ACM\/IEEE International Conference on Human-Robot Interaction (HRI 2010), Osaka, Japan.","DOI":"10.1109\/HRI.2010.5453183"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Aran, O., and Gatica-Perez, D. (2013, January 9\u201313). One of a Kind: Inferring Personality Impressions in Meetings. Proceedings of the 15th ACM International Conference on Multimodal Interaction (ICMI 2013), Sydney, Australia.","DOI":"10.1145\/2522848.2522859"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Okada, S., Nakano, Y., Hayashi, Y., Takase, Y., and Nitta, K. (2016, January 12\u201316). Estimating Communication Skills using Dialogue Acts and Nonverbal Features in Multiple Discussion Datasets. Proceedings of the 18th ACM International Conference on Multimodal Interaction (ICMI 2016), Tokyo, Japan.","DOI":"10.1145\/2993148.2993154"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Schiavo, G., Cappelletti, A., Mencarini, E., Stock, O., and Zancanaro, M. (2014, January 24\u201327). Overt or Subtle? Supporting Group Conversations with Automatically Targeted Directives. Proceedings of the 19th international conference on Intelligent User Interfaces (IUI 2014), Haifa, Israel.","DOI":"10.1145\/2557500.2557507"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Raducanu, B., Vitria, J., and Gatica-Perez, D. (2009, January 19\u201324). You are fired! Nonverbal role analysis in competitive meetings. Proceedings of the 2009 IEEE International Conference onAcoustics, Speech and Signal Processing (ICASSP 2009), Taipei, Taiwan.","DOI":"10.1109\/ICASSP.2009.4959992"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Muralidhar, S., Nguyen, L.S., Frauendorfer, D., Odobez, J.M., Mast, M.S., and Gatica-Perez, D. (2016, January 12\u201316). Training on the Job: Behavioral Analysis of Job Interviews in Hospitality. Proceedings of the 18th ACM International Conference on Multimodal Interaction (ICMI 2016), Tokyo, Japan.","DOI":"10.1145\/2993148.2993191"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Oertel, C., Mora, K.A.F., Gustafson, J., and Odobez, J.M. (2015, January 9\u201313). Deciphering the Silent Participant: On the Use of Audio-Visual Cues for the Classification of Listener Categories in Group Discussions. Proceedings of the 17th ACM on International Conference on Multimodal Interaction (ICMI 2015), Seattle, WA, USA.","DOI":"10.1145\/2818346.2820759"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Bevacqua, E., Pammi, S., Hyniewska, S., Schroder, M., and Pelachaud, C. (2010, January 20\u201322). Multimodal Backchannels for Embodied Conversational Agents. Proceedings of the 10th International Conference on Intelligent Virtual Agents (IVA 2010), Philadelphia, PA, USA.","DOI":"10.1007\/978-3-642-15892-6_21"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Oertel, C., Lopes, J., Yu, Y., Mora, K.A.F., Gustafson, J., Black, A.W., and Odobez, J.M. (2016, January 12\u201316). Towards Building an Attentive Artificial Listener: On the Perception of Attentiveness in Audio-Visual Feedback Tokens. Proceedings of the 18th ACM International Conference on Multimodal Interaction (ICMI 2016), Tokyo, Japan.","DOI":"10.1145\/2993148.2993188"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Agarwal, P., Moubayed, S.A., Alspach, A., Kim, J., Carter, E.J., Lehman, J., and Yamane, K. (2016, January 26\u201331). Imitating human movement with teleoperated robotic head. Proceedings of the 25th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN 2016), New York, NY, USA.","DOI":"10.1109\/ROMAN.2016.7745184"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Cazzato, D., Cimarelli, C., Sanchez-Lopez, J.L., Olivares-Mendez, M.A., and Voos, H. (2019, January 27\u201329). Real-Time Human Head Imitation for Humanoid Robots. Proceedings of the 3rd International Conference on Artificial Intelligence and Virtual Reality (AIVR 2019), Singapore.","DOI":"10.1145\/3348488.3348501"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Ondras, J., Celiktutan, O., Sariyanidi, E., and Gunes, H. (2017, January 28\u201331). Automatic replication of teleoperator head movements and facial expressions on a humanoid robot. Proceedings of the 26th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN 2017), Lisbon, Portugal.","DOI":"10.1109\/ROMAN.2017.8172386"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"25","DOI":"10.5898\/JHRI.6.1.Admoni","article-title":"Social eye gaze in human-robot interaction: A review","volume":"6","author":"Admoni","year":"2017","journal-title":"J. Hum.-Robot Interact."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Leite, I., McCoy, M., Lohani, M., Ullman, D., Salomons, N., Stokes, C., Rivers, S., and Scassellati, B. (2015, January 2\u20135). Emotional Storytelling in the Classroom: Individual versus Group Interaction between Children and Robots. Proceedings of the 10th ACM\/IEEE International Conference on Human-Robot Interaction (HRI 2015), Portland, OR, USA.","DOI":"10.1145\/2696454.2696481"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Vazquez, M., Carter, E.J., McDorman, B., Steinfeld, J.F.A., and Hudson, S.E. (2017, January 6\u20139). Towards Robot Autonomy in Group Conversations: Understanding the Effects of Body Orientation and Gaze. Proceedings of the 12th ACM\/IEEE International Conference on Human-Robot Interaction (HRI 2017), Vienna, Austria.","DOI":"10.1145\/2909824.3020207"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"557","DOI":"10.3389\/fpsyg.2013.00557","article-title":"Automatic detection of service initiation signals used in bars","volume":"4","author":"Loth","year":"2013","journal-title":"Front. Psychol."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Keizer, S., Foster, M.E., Wang, Z., and Lemon, O.J. (2014). Machine Learning for Social Multi-Party Human-Robot Interaction. ACM Trans. Interact. Intell. Syst., 4.","DOI":"10.1145\/2600021"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"686","DOI":"10.20965\/jaciii.2017.p0686","article-title":"Generation of Bystander Robot Actions Based on Analysis of Relative Probability of Human Actions","volume":"21","author":"Sakai","year":"2017","journal-title":"J. Adv. Comput. Intell. Intell. Inform."},{"key":"ref_38","unstructured":"Bohus, D., and Horvitz, E. (2014, January 12\u201316). Managing Human-Robot Engagement with Forecasts and... Um... Hesitations. Proceedings of the 16th International Conference on Multimodal Interaction (ICMI 2014), Istanbul, Turkey."},{"key":"ref_39","unstructured":"Sidner, C.L., Lee, C., and Lesh, N. (2003). Engagement Rules for Human-Robot Collaborative Interaction, Mitsubishi Electric Research Laboratories. Technical Report TR2003-50."},{"key":"ref_40","first-page":"12:1","article-title":"Conversational Gaze Mechanisms for Humanlike Robots","volume":"1","author":"Mutlu","year":"2012","journal-title":"ACM Trans. Interact. Intell. Syst. (TiiS)"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Bohus, D., and Horvitz, E. (2009, January 11\u201312). Learning to Predict Engagement with a Spoken Dialog System in Open-World Settings. Proceedings of the 10th Annual Meeting of the Special Interest Group on Discourse and Dialogue, London, UK.","DOI":"10.3115\/1708376.1708411"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Sakai, K., Ishi, C.T., Minato, T., and Ishiguro, H. (September, January 31). Online speech-driven head motion generating system and evaluation on a tele-operated robot. Proceedings of the 24th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN 2015), Kobe, Japan.","DOI":"10.1109\/ROMAN.2015.7333610"},{"key":"ref_43","first-page":"8","article-title":"Modeling of Human Visual Attention in Multiparty Open-World Dialogues","volume":"8","author":"Stefanov","year":"2019","journal-title":"ACM Trans. Hum.-Robot Interact. (THRI)"},{"key":"ref_44","unstructured":"Stefanov, K., and Beskow, J. (2016, January 23\u201328). A Multi-party Multi-modal Dataset for Focus of Visual Attention in Human-human and Human-robot Interaction. Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016), Portoroz, Slovenia."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Nihei, F., Nakano, Y.I., Hayashi, Y., Huang, H.H., and Okada, S. (2014, January 12\u201316). Predicting Influential Statements in Group Discussions using Speech and Head Motion Information. Proceedings of the 16th International Conference on Multimodal Interaction (ICMI 2014), Istanbul, Turkey.","DOI":"10.1145\/2663204.2663248"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1023\/A:1007714801558","article-title":"Emergent Leadership Behaviors: The Function of Personality and Cognitive Ability in Determining Teamwork Performance and KSAs","volume":"15","author":"Kickul","year":"2000","journal-title":"J. Bus. Psychol."},{"key":"ref_47","unstructured":"Boersma, P., and Weenink, D. (2019, May 26). Praat: Doing Phonetics by Computer [Computer Software] Version 6.0.40. Available online: http:\/\/www.praat.org\/."},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"841","DOI":"10.3758\/BRM.41.3.841","article-title":"Coding gestural behavior with the NEUROGES\u2013ELAN system","volume":"41","author":"Lausberg","year":"2009","journal-title":"Behav. Res. Methods"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1613\/jair.953","article-title":"SMOTE: Synthetic Minority Over-sampling Technique","volume":"16","author":"Chawla","year":"2002","journal-title":"J. Artif. Intell. Res."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Morikawa, O., and Maesako, T. (1998, January 14\u201318). HyperMirror: Toward Pleasant-to-use Video Mediated Communication System. Proceedings of the 1998 ACM Conference on Computer Supported Cooperative Work (CSCW\u201998), Seattle, WA, USA.","DOI":"10.1145\/289444.289489"}],"container-title":["Multimodal Technologies and Interaction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2414-4088\/4\/2\/15\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T14:09:26Z","timestamp":1760364566000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2414-4088\/4\/2\/15"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,4,29]]},"references-count":50,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2020,6]]}},"alternative-id":["mti4020015"],"URL":"https:\/\/doi.org\/10.3390\/mti4020015","relation":{},"ISSN":["2414-4088"],"issn-type":[{"value":"2414-4088","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,4,29]]}}}