{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T17:01:44Z","timestamp":1784221304685,"version":"3.55.0"},"reference-count":38,"publisher":"Wiley","issue":"1","license":[{"start":{"date-parts":[[2023,12,23]],"date-time":"2023-12-23T00:00:00Z","timestamp":1703289600000},"content-version":"vor","delay-in-days":356,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2022YFF0902701"],"award-info":[{"award-number":["2022YFF0902701"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62202065"],"award-info":[{"award-number":["62202065"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U21A20468"],"award-info":[{"award-number":["U21A20468"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61921003"],"award-info":[{"award-number":["61921003"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61972043"],"award-info":[{"award-number":["61972043"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"publisher","award":["2020XD-A07-1"],"award-info":[{"award-number":["2020XD-A07-1"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["International Journal of Intelligent Systems"],"published-print":{"date-parts":[[2023,1]]},"abstract":"<jats:p>Intelligent service robots have become an indispensable aspect of modern\u2010day society, playing a crucial role in various domains ranging from healthcare to hospitality. Among these robotic systems, human\u2010machine dialogue systems are particularly noteworthy as they deliver both auditory and visual services to users, effectively bridging the communication gap between humans and machines. Despite their utility, the majority of existing approaches to these systems primarily concentrate on augmenting the logical coherence of the system\u2019s responses, inadvertently neglecting the significance of user emotions in shaping a comprehensive communication experience. To tackle this shortcoming, we propose the development of an innovative human\u2010machine dialogue system that is both intelligent and emotionally sensitive, employing multimodal generation techniques. This system is architecturally comprised of three components: (1) data collection and processing, responsible for gathering and preparing relevant information, (2) a dialogue engine, which generates contextually appropriate responses, and (3) an interaction module, responsible for facilitating the communication interface between users and the system. To validate our proposed approach, we have constructed a prototype system and conducted an evaluation of the performance of the core dialogue engine by utilizing an open dataset. The results of our study indicate that our system demonstrates a remarkable level of multimodal generation response, ultimately offering a more human\u2010like dialogue experience.<\/jats:p>","DOI":"10.1155\/2023\/9267487","type":"journal-article","created":{"date-parts":[[2023,12,23]],"date-time":"2023-12-23T18:20:05Z","timestamp":1703355605000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":13,"title":["Beyond Words: An Intelligent Human\u2010Machine Dialogue System with Multimodal Generation and Emotional Comprehension"],"prefix":"10.1155","volume":"2023","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2151-5420","authenticated-orcid":false,"given":"Yaru","family":"Zhao","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2160-2839","authenticated-orcid":false,"given":"Bo","family":"Cheng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yakun","family":"Huang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhiguo","family":"Wan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2023,12,23]]},"reference":[{"key":"e_1_2_10_1_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2022.108318"},{"key":"e_1_2_10_2_2","unstructured":"SutskeverI. VinyalsO. andQuocV. L. Sequence to sequence learning with neural networks Proceedings of the 27th International Conference on Neural Information Processing Systems December 2014 Montreal Quebec Canada 3104\u20133112."},{"key":"e_1_2_10_3_2","unstructured":"VaswaniA. ShazeerN. andParmarN. Attention is all you need Proceedings of the 31st International Conference on Neural Information Processing Systems December 2017 Long Beach CA USA 6000\u20136010."},{"key":"e_1_2_10_4_2","doi-asserted-by":"crossref","unstructured":"ShusterK. HumeauS. AntoineB. andWestonJ. Image chat: engaging grounded conversations Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics July 2020 Stroudsburg PA USA 2414\u20132429.","DOI":"10.18653\/v1\/2020.acl-main.219"},{"key":"e_1_2_10_5_2","doi-asserted-by":"crossref","unstructured":"ShusterK. Michael SmithE. JuD. andWestonJ. Multi-modal open-domain dialogue Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing November 2021 Punta Cana Dominican Republic 4863\u20134883.","DOI":"10.18653\/v1\/2021.emnlp-main.398"},{"key":"e_1_2_10_6_2","doi-asserted-by":"crossref","unstructured":"WeiW. LiuJ. andMaoX. Emotion-aware chat machine: automatic emotional response generation for human-like emotional interaction Proceedings of the 28th ACM International Conference on Information and Knowledge Management November 2019 Beijing China 1401\u20131410.","DOI":"10.1145\/3357384.3357937"},{"key":"e_1_2_10_7_2","doi-asserted-by":"crossref","unstructured":"LiS. ShiF. andWangD. Emoelicitor: an open domain response generation model with user emotional reaction awareness Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence January 2021 Yokohama Japan 3637\u20133643.","DOI":"10.24963\/ijcai.2020\/503"},{"key":"e_1_2_10_8_2","unstructured":"FungP. DeyA. andSiddiqueF. B. Zara: a virtual interactive dialogue system incorporating emotion sentiment and personality recognition Proceedings of COLING 2016 the 26th International Conference on Computational Linguistics December 2016 Yokohama Japan System Demonstrations 278\u2013281."},{"key":"e_1_2_10_9_2","doi-asserted-by":"crossref","unstructured":"HuberB. McDuffD. BrockettC. GalleyM. andDolanB. Emotional dialogue generation using image-grounded language models Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems April 2018 Montreal QC Canada 1\u201312.","DOI":"10.1145\/3173574.3173851"},{"key":"e_1_2_10_10_2","doi-asserted-by":"crossref","unstructured":"TianZ. WenZ. WuZ. SongY. andTangJ. Emotion-aware multimodal pre-training for image-grounded emotional response generation Proceedings of the International Conference on Database Systems for Advanced Applications April 2022 Tianjin China 3\u201319.","DOI":"10.1007\/978-3-031-00129-1_1"},{"key":"e_1_2_10_11_2","doi-asserted-by":"crossref","unstructured":"ShenT. ZuoJ. FanS. ZhangJ. andJiangL. Vida-man: visual dialog with digital humans Proceedings of the 29th ACM International Conference on Multimedia October 2021 Verlagsort NY USA 2789\u20132791.","DOI":"10.1145\/3474085.3478560"},{"key":"e_1_2_10_12_2","unstructured":"WangS. MengY. andSunX. Modeling text-visual mutual dependency for multi-modal dialog generation 2021 https:\/\/arxiv.org\/abs\/2105.14445."},{"key":"e_1_2_10_13_2","doi-asserted-by":"crossref","unstructured":"CastellanoG. De CarolisB. MarvulliN. SciancaleporeM. andVessioG. Real-time age estimation from facial images using yolo and efficientnet International Conference on Computer Analysis of Images and Patterns September 2021 Limassol Cyprus 275\u2013284.","DOI":"10.1007\/978-3-030-89131-2_25"},{"key":"e_1_2_10_14_2","unstructured":"DevlinJ. ChangM.-W. LeeK. andToutanovaK. Bert: pre-training of deep bidirectional transformers for language understanding Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies October2019 Stroudsburg PA USA 4171\u20134186."},{"key":"e_1_2_10_15_2","doi-asserted-by":"crossref","unstructured":"XieS. GirshickR. Doll\u00e1rP. TuZ. andHeK. Aggregated residual transformations for deep neural networks Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition June 2017 Las Vegas ND USA 1492\u20131500.","DOI":"10.1109\/CVPR.2017.634"},{"key":"e_1_2_10_16_2","doi-asserted-by":"crossref","unstructured":"ZhouH. YoungT. HuangM.et al. Commonsense knowledge aware conversation generation with graph attention Proceedings of the 27th International Joint Conference on Artificial Intelligence July2018 Stockholm Sweden 4623\u20134629.","DOI":"10.24963\/ijcai.2018\/643"},{"key":"e_1_2_10_17_2","unstructured":"XiongC. ZhongV. andSocherR. Dynamic coattention networks for question answering Proceedings of the 5th International Conference on Learning Representations April 2017 Toulon France."},{"key":"e_1_2_10_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_2_10_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4842-5364-9_2"},{"key":"e_1_2_10_20_2","unstructured":"ChenY. WuL. andZakiM. J. Iterative deep graph learning for graph neural networks: better and robust node embeddings Proceedings of the 34th International Conference on Neural Information Processing Systems December 2020 Canada 19314\u201319326."},{"key":"e_1_2_10_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/89.966080"},{"key":"e_1_2_10_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10579-008-9076-6"},{"key":"e_1_2_10_23_2","doi-asserted-by":"crossref","unstructured":"SpeerR. ChinJ. andHavasiC. Conceptnet 5.5: an open multilingual graph of general knowledge Proceedings of the 31st AAAI Conference on Artificial Intelligence February 2017 San Francisco CA USA 4444\u20134451.","DOI":"10.1609\/aaai.v31i1.11164"},{"key":"e_1_2_10_24_2","doi-asserted-by":"crossref","unstructured":"GuanJ. WangY. andHuangM. Story ending generation with incremental encoding and commonsense knowledge Proceedings of the 33rd AAAI Conference on Artificial Intelligence February 2019 Honolulu HW USA 6473\u20136480.","DOI":"10.1609\/aaai.v33i01.33016473"},{"key":"e_1_2_10_25_2","doi-asserted-by":"crossref","unstructured":"SongH. WangY. ZhangW. LiuX. andLiuT. Generate delete and rewrite: a three-stage framework for improving persona consistency of dialogue generation Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics July 2020 Stroudsburg PA USA 5821\u20135831.","DOI":"10.18653\/v1\/2020.acl-main.516"},{"key":"e_1_2_10_26_2","doi-asserted-by":"crossref","unstructured":"LiJ. MonroeW. andJurafskyD. A diversity-promoting objective function for neural conversation models Proceedings of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies July 2016 Stroudsburg PA USA 110\u2013119.","DOI":"10.18653\/v1\/N16-1014"},{"key":"e_1_2_10_27_2","doi-asserted-by":"crossref","unstructured":"GeorgeD. Automatic evaluation of machine translation quality using n-gram co-occurrence statistics Proceedings of the Second International Conference on Human Language Technology Research November 2002 Stroudsburg PA USA 138\u2013145.","DOI":"10.3115\/1289189.1289273"},{"key":"e_1_2_10_28_2","first-page":"74","volume-title":"Proceeding of the Workshop on Text Summariation Branches Out, Post-Conference Workshop of ACL 2004","author":"Lin C.-Y.","year":"2004"},{"key":"e_1_2_10_29_2","unstructured":"BanerjeeS.andAlonL. Meteor: an automatic metric for mt evaluation with improved correlation with human judgments Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and\/or summarization June 2005 Ann Arbor MG USA 65\u201372."},{"key":"e_1_2_10_30_2","doi-asserted-by":"crossref","unstructured":"LiC.-Y. OrtegaD. andV\u00e4thD. Adviser: a toolkit for developing multi-modal multi-domain and socially-engaged conversational agents Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations July 2020 Stroudsburg PA USA 279\u2013286.","DOI":"10.18653\/v1\/2020.acl-demos.31"},{"key":"e_1_2_10_31_2","unstructured":"MouL. SongY. YanR. LiG. andZhangL. Sequence to backward and forward sequences: a content-introducing approach to generative short-text conversation Proceedings of the 26th International Conference on Computational Linguistics December 2016 Osaka Japan 3349\u20133358."},{"key":"e_1_2_10_32_2","doi-asserted-by":"crossref","unstructured":"PangB. NijkampE. andHanW. Towards holistic and automatic evaluation of open-domain dialogue generation Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics July 2020 Stroudsburg PA USA 3619\u20133629.","DOI":"10.18653\/v1\/2020.acl-main.333"},{"key":"e_1_2_10_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/t-affc.2013.11"},{"key":"e_1_2_10_34_2","unstructured":"SanoY. LeowC. S. andIidayS. Spoken dialog training system for customer service improvement Proceedings of the 2020 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference December 2020 Auckland New Zealand 403\u2013408."},{"key":"e_1_2_10_35_2","unstructured":"ChenJ. SunJ. andHuangH. An open-source dialog system with real-time engagement tracking for job interview training applications Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations December 2020 Berlin Germany 10\u201315."},{"key":"e_1_2_10_36_2","doi-asserted-by":"crossref","unstructured":"CuiC. WangW. SongX. HuangM. andXuX. S. User attention-guided multimodal dialog systems Proceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval July 2019 Paris France 445\u2013454.","DOI":"10.1145\/3331184.3331226"},{"key":"e_1_2_10_37_2","doi-asserted-by":"crossref","unstructured":"LiaoL. MaY. HeX. HongR. andChuaT.-S. Knowledge-aware multimodal dialogue systems Proceedings of the 26th ACM International Conference on Multimedia October 2018 Seoul Republic of Korea 801\u2013809.","DOI":"10.1145\/3240508.3240605"},{"key":"e_1_2_10_38_2","doi-asserted-by":"crossref","unstructured":"SunQ. WangY. andXuC. Multimodal dialogue response generation Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics May 2022 Dublin Ireland 2854\u20132866.","DOI":"10.18653\/v1\/2022.acl-long.204"}],"container-title":["International Journal of Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/downloads.hindawi.com\/journals\/ijis\/2023\/9267487.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/downloads.hindawi.com\/journals\/ijis\/2023\/9267487.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1155\/2023\/9267487","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,31]],"date-time":"2024-12-31T05:34:05Z","timestamp":1735623245000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1155\/2023\/9267487"}},"subtitle":[],"editor":[{"given":"Alexander","family":"Ho\u0161ovsk\u00fd","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2023,1]]},"references-count":38,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2023,1]]}},"alternative-id":["10.1155\/2023\/9267487"],"URL":"https:\/\/doi.org\/10.1155\/2023\/9267487","archive":["Portico"],"relation":{},"ISSN":["0884-8173","1098-111X"],"issn-type":[{"value":"0884-8173","type":"print"},{"value":"1098-111X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1]]},"assertion":[{"value":"2023-04-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-12-23","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"9267487"}}