{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,16]],"date-time":"2026-04-16T03:38:01Z","timestamp":1776310681855,"version":"3.50.1"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,11,7]],"date-time":"2023-11-07T00:00:00Z","timestamp":1699315200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Project of New Generation Artificial Intelligence","award":["2018AAA0102502"],"award-info":[{"award-number":["2018AAA0102502"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U1936203"],"award-info":[{"award-number":["U1936203"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100007129","name":"Shandong Provincial Natural Science Foundation","doi-asserted-by":"crossref","award":["ZR2022YQ59"],"award-info":[{"award-number":["ZR2022YQ59"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2024,3,31]]},"abstract":"<jats:p>\n            Text response generation for multimodal task-oriented dialog systems, which aims to generate the proper text response given the multimodal context, is an essential yet challenging task. Although existing efforts have achieved compelling success, they still suffer from two pivotal limitations: (1)\u00a0\n            <jats:italic>overlook the benefit of generative pretraining<\/jats:italic>\n            and (2)\n            <jats:italic>ignore the textual context-related knowledge<\/jats:italic>\n            . To address these limitations, we propose a novel dual knowledge-enhanced generative pretrained language mode for multimodal task-oriented dialog systems\u00a0(DKMD), consisting of three key components:\n            <jats:italic>dual knowledge selection<\/jats:italic>\n            ,\n            <jats:italic>dual knowledge-enhanced context learning<\/jats:italic>\n            , and\n            <jats:italic>knowledge-enhanced response generation<\/jats:italic>\n            . To be specific, the dual knowledge selection component aims to select the related knowledge according to both textual and visual modalities of the given context. Thereafter, the dual knowledge-enhanced context learning component targets seamlessly, integrating the selected knowledge into the multimodal context learning from both global and local perspectives, where the cross-modal semantic relation is also explored. Moreover, the knowledge-enhanced response generation component comprises a revised BART decoder, where an additional dot-product knowledge-decoder attention sub-layer is introduced for explicitly utilizing the knowledge to advance the text response generation. Extensive experiments on a public dataset verify the superiority of the proposed DKMD over state-of-the-art competitors.\n          <\/jats:p>","DOI":"10.1145\/3606368","type":"journal-article","created":{"date-parts":[[2023,10,6]],"date-time":"2023-10-06T15:45:55Z","timestamp":1696607155000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":19,"title":["Multimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model"],"prefix":"10.1145","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4638-0603","authenticated-orcid":false,"given":"Xiaolin","family":"Chen","sequence":"first","affiliation":[{"name":"School of Software, Joint SDU-NTU Centre for Artificial Intelligence Research, Shandong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5274-4197","authenticated-orcid":false,"given":"Xuemeng","family":"Song","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Shandong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9827-5835","authenticated-orcid":false,"given":"Liqiang","family":"Jing","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Shandong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-2204-2066","authenticated-orcid":false,"given":"Shuo","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Shandong University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0993-8495","authenticated-orcid":false,"given":"Linmei","family":"Hu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Beijing Institute of Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1476-0273","authenticated-orcid":false,"given":"Liqiang","family":"Nie","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,11,7]]},"reference":[{"issue":"1","key":"e_1_3_2_2_2","first-page":"5:1\u20135:24","article-title":"CATS: Customizable abstractive topic-based summarization","volume":"40","author":"Bahrainian Seyed Ali","year":"2022","unstructured":"Seyed Ali Bahrainian, George Zerveas, Fabio Crestani, and Carsten Eickhoff. 2022. CATS: Customizable abstractive topic-based summarization. ACM Trans. Inf. Syst. 40, 1 (2022), 5:1\u20135:24.","journal-title":"ACM Trans. Inf. Syst."},{"key":"e_1_3_2_3_2","first-page":"1665","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Cai Jie","year":"2020","unstructured":"Jie Cai, Zhengzhou Zhu, Ping Nie, and Qian Liu. 2020. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 1665\u20131668."},{"key":"e_1_3_2_4_2","first-page":"2819","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Chaplot Devendra Singh","year":"2018","unstructured":"Devendra Singh Chaplot, Kanthashree Mysore Sathyendra, Rama Kumar Pasumarthi, Dheeraj Rajagopal, and Ruslan Salakhutdinov. 2018. Gated-attention architectures for task-oriented language grounding. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 2819\u20132826."},{"key":"e_1_3_2_5_2","first-page":"1466","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Chatterjee Shubham","year":"2022","unstructured":"Shubham Chatterjee and Laura Dietz. 2022. BERT-ER: Query-specific BERT entity representations for entity ranking. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 1466\u20131477."},{"key":"e_1_3_2_6_2","first-page":"5437","volume-title":"Proceedings of the Conference of the Association for Computational Linguistics","author":"Chauhan Hardik","year":"2019","unstructured":"Hardik Chauhan, Mauajama Firdaus, Asif Ekbal, and Pushpak Bhattacharyya. 2019. Ordinal and attribute aware response generation in a multimodal dialogue system. In Proceedings of the Conference of the Association for Computational Linguistics. Association for Computational Linguistics, 5437\u20135447."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2020.3014166"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462946"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3406109"},{"key":"e_1_3_2_10_2","unstructured":"Junyoung Chung \u00c7aglar G\u00fcl\u00e7ehre KyungHyun Cho and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. CoRR abs\/1412.3555 (2014). Retrieved from http:\/\/arxiv.org\/abs\/1412.3555"},{"key":"e_1_3_2_11_2","first-page":"445","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Cui Chen","year":"2019","unstructured":"Chen Cui, Wenjie Wang, Xuemeng Song, Minlie Huang, Xin-Shun Xu, and Liqiang Nie. 2019. User attention-guided multimodal dialog systems. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 445\u2013454."},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3507782"},{"key":"e_1_3_2_13_2","first-page":"4171","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics. Association for Computational Linguistics, 4171\u20134186."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.5555\/1289189.1289273"},{"key":"e_1_3_2_15_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In Proceedings of the International Conference on Learning Representations. OpenReview.net."},{"issue":"4","key":"e_1_3_2_16_2","first-page":"68:1\u201368:14","article-title":"Text-based editing of talking-head video","volume":"38","author":"Fried Ohad","year":"2019","unstructured":"Ohad Fried, Ayush Tewari, Michael Zollh\u00f6fer, Adam Finkelstein, Eli Shechtman, Dan B. Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, and Maneesh Agrawala. 2019. Text-based editing of talking-head video. ACM Trans. Graph. 38, 4 (2019), 68:1\u201368:14.","journal-title":"ACM Trans. Graph."},{"issue":"4","key":"e_1_3_2_17_2","first-page":"81:1\u201381:32","article-title":"\u201cWhat can I cook with these ingredients?\u201d\u2014Understanding cooking-related information needs in conversational search","volume":"40","author":"Frummet Alexander","year":"2022","unstructured":"Alexander Frummet, David Elsweiler, and Bernd Ludwig. 2022. \u201cWhat can I cook with these ingredients?\u201d\u2014Understanding cooking-related information needs in conversational search. ACM Trans. Inf. Syst. 40, 4 (2022), 81:1\u201381:32.","journal-title":"ACM Trans. Inf. Syst."},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01026"},{"key":"e_1_3_2_19_2","first-page":"307","volume-title":"Proceedings of the ACM International Conference on Web Search and Data Mining","author":"Gao Shen","year":"2022","unstructured":"Shen Gao, Yuchi Zhang, Yongliang Wang, Yang Dong, Xiuying Chen, Dongyan Zhao, and Rui Yan. 2022. HeteroQA: Learning towards question-and-answering through multiple information sources via heterogeneous graph modeling. In Proceedings of the ACM International Conference on Web Search and Data Mining. ACM, 307\u2013315."},{"key":"e_1_3_2_20_2","first-page":"2755","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"He Weidong","year":"2020","unstructured":"Weidong He, Zhi Li, Dongcai Lu, Enhong Chen, Tong Xu, Baoxing Huai, and Jing Yuan. 2020. Multimodal dialogue systems via capturing context-aware dependencies of semantic elements. In Proceedings of the ACM International Conference on Multimedia. ACM, 2755\u20132764."},{"key":"e_1_3_2_21_2","first-page":"1657","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Kang Taegwan","year":"2021","unstructured":"Taegwan Kang, Hwanhee Lee, Byeongjin Choe, and Kyomin Jung. 2021. Entangled bidirectional encoder to autoregressive decoder for sequential recommendation. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 1657\u20131661."},{"key":"e_1_3_2_22_2","first-page":"28","volume-title":"Proceedings of the 11th ACM International Conference on Web Search and Data Mining","author":"Kordan Saeid Balaneshin","year":"2018","unstructured":"Saeid Balaneshin Kordan and Alexander Kotov. 2018. Deep neural architecture for multi-modal retrieval based on joint embedding space for text and images. In Proceedings of the 11th ACM International Conference on Web Search and Data Mining. ACM, 28\u201336."},{"key":"e_1_3_2_23_2","first-page":"1437","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Lei Wenqiang","year":"2018","unstructured":"Wenqiang Lei, Xisen Jin, Min-Yen Kan, Zhaochun Ren, Xiangnan He, and Dawei Yin. 2018. Sequicity: Simplifying task-oriented dialogue systems with single sequence-to-sequence architectures. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 1437\u20131447."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1016\/0031-3203(93)90115-D"},{"key":"e_1_3_2_26_2","first-page":"11002","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Li Shimin","year":"2022","unstructured":"Shimin Li, Hang Yan, and Xipeng Qiu. 2022. Contrast and generation make BART a good dialogue emotion recognizer. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 11002\u201311010."},{"key":"e_1_3_2_27_2","first-page":"733","volume-title":"Proceedings of the 8th International Joint Conference on Natural Language Processing","author":"Li Xiujun","year":"2017","unstructured":"Xiujun Li, Yun-Nung Chen, Lihong Li, Jianfeng Gao, and Asli Celikyilmaz. 2017. End-to-end task-completion neural dialogue systems. In Proceedings of the 8th International Joint Conference on Natural Language Processing. Asian Federation of Natural Language Processing, 733\u2013743."},{"key":"e_1_3_2_28_2","first-page":"1930","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Li Yanran","year":"2021","unstructured":"Yanran Li, Wenjie Li, and Zhitao Wang. 2021. Graph-structured context understanding for knowledge-grounded response generation. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 1930\u20131934."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2020.2977637"},{"key":"e_1_3_2_30_2","first-page":"675","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Liao Lizi","year":"2021","unstructured":"Lizi Liao, Le Hong Long, Zheng Zhang, Minlie Huang, and Tat-Seng Chua. 2021. MMConv: An environment for multimodal conversational search across multiple domains. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 675\u2013684."},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"801","DOI":"10.1145\/3240508.3240605","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Liao Lizi","year":"2018","unstructured":"Lizi Liao, Yunshan Ma, Xiangnan He, Richang Hong, and Tat-Seng Chua. 2018. Knowledge-aware multimodal dialogue systems. In Proceedings of the ACM International Conference on Multimedia. ACM, 801\u2013809."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2020.3008563"},{"key":"e_1_3_2_33_2","first-page":"1","article-title":"Graph-grounded goal planning for conversational recommendation","author":"Liu Zeming","year":"2022","unstructured":"Zeming Liu, Ding Zhou, Hao Liu, Haifeng Wang, Zheng-Yu Niu, Hua Wu, Wanxiang Che, Ting Liu, and Hui Xiong. 2022. Graph-grounded goal planning for conversational recommendation. IEEE Trans. Knowl. Data Eng. 35, 5 (2022), 1\u201315.","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"e_1_3_2_34_2","first-page":"103","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Ma Zhiyuan","year":"2022","unstructured":"Zhiyuan Ma, Jianjun Li, Guohui Li, and Yongjing Cheng. 2022. UniTranSeR: A unified transformer semantic representation framework for multimodal task-oriented dialog system. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 103\u2013114."},{"key":"e_1_3_2_35_2","first-page":"1468","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Madotto Andrea","year":"2018","unstructured":"Andrea Madotto, Chien-Sheng Wu, and Pascale Fung. 2018. Mem2Seq: Effectively incorporating knowledge bases into end-to-end task-oriented dialog systems. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 1468\u20131478."},{"key":"e_1_3_2_36_2","first-page":"3111","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In Proceedings of the International Conference on Neural Information Processing Systems. Curran Associates Inc., 3111\u20133119."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3470562"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3108724"},{"key":"e_1_3_2_39_2","first-page":"1098","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Nie Liqiang","year":"2019","unstructured":"Liqiang Nie, Wenjie Wang, Richang Hong, Meng Wang, and Qi Tian. 2019. Multimodal dialog system: Generating responses via adaptive decoders. In Proceedings of the ACM International Conference on Multimedia. ACM, 1098\u20131106."},{"key":"e_1_3_2_40_2","first-page":"311","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. ACL, 311\u2013318."},{"key":"e_1_3_2_41_2","first-page":"8024","volume-title":"Proceedings of the Annual Conference on Neural Information Processing Systems","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K\u00f6pf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of the Annual Conference on Neural Information Processing Systems. 8024\u20138035."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_43_2","first-page":"8748","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_2_44_2","unstructured":"Alec Radford and Karthik Narasimhan. 2018. Improving language understanding by generative pre-training."},{"key":"e_1_3_2_45_2","first-page":"140:1\u2013140:67","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21 (2020), 140:1\u2013140:67.","journal-title":"J. Mach. Learn. Res."},{"issue":"4","key":"e_1_3_2_46_2","first-page":"47:1\u201347:29","article-title":"Conversations with search engines: SERP-based conversational response generation","volume":"39","author":"Ren Pengjie","year":"2021","unstructured":"Pengjie Ren, Zhumin Chen, Zhaochun Ren, Evangelos Kanoulas, Christof Monz, and Maarten de Rijke. 2021. Conversations with search engines: SERP-based conversational response generation. ACM Trans. Inf. Syst. 39, 4 (2021), 47:1\u201347:29.","journal-title":"ACM Trans. Inf. Syst."},{"key":"e_1_3_2_47_2","first-page":"696","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Saha Amrita","year":"2018","unstructured":"Amrita Saha, Mitesh M. Khapra, and Karthik Sankaranarayanan. 2018. Towards building large scale multimodal domain-aware conversation systems. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 696\u2013704."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01740"},{"key":"e_1_3_2_49_2","first-page":"992","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Song Xuemeng","year":"2022","unstructured":"Xuemeng Song, Liqiang Jing, Dengtian Lin, Zhongzhou Zhao, Haiqing Chen, and Liqiang Nie. 2022. V2P: Vision-to-prompt based multi-modal product summary generation. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 992\u20131001."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01813"},{"key":"e_1_3_2_51_2","first-page":"3104","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems","author":"Sutskever Ilya","year":"2014","unstructured":"Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In Proceedings of the International Conference on Neural Information Processing Systems. MIT Press, 3104\u20133112."},{"key":"e_1_3_2_52_2","first-page":"5998","volume-title":"Proceedings of the International Conference on Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the International Conference on Advances in Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_2_53_2","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N. Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017 December 4-9 2017 Long Beach CA 5998\u20136008. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html"},{"key":"e_1_3_2_54_2","first-page":"4100","volume-title":"Proceedings of the International Conference on Computational Linguistics","author":"Wang Jian","year":"2020","unstructured":"Jian Wang, Junhao Liu, Wei Bi, Xiaojiang Liu, Kejing He, Ruifeng Xu, and Min Yang. 2020. Dual dynamic memory network for end-to-end multi-turn task-oriented dialog systems. In Proceedings of the International Conference on Computational Linguistics. International Committee on Computational Linguistics, 4100\u20134110."},{"key":"e_1_3_2_55_2","first-page":"1369","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Wen Haokun","year":"2021","unstructured":"Haokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan, and Liqiang Nie. 2021. Comprehensive linguistic-visual composition network for image retrieval. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 1369\u20131378."},{"key":"e_1_3_2_56_2","first-page":"1901","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Xu Weidi","year":"2020","unstructured":"Weidi Xu, Xingyi Cheng, Kunlong Chen, and Taifeng Wang. 2020. Symmetric regularization based BERT for pair-wise semantic reasoning. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 1901\u20131904."},{"key":"e_1_3_2_57_2","first-page":"4970","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Young Tom","year":"2018","unstructured":"Tom Young, Erik Cambria, Iti Chaturvedi, Hao Zhou, Subham Biswas, and Minlie Huang. 2018. Augmenting end-to-end dialogue systems with commonsense knowledge. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 4970\u20134977."},{"key":"e_1_3_2_58_2","first-page":"3995","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Yu Tiezheng","year":"2021","unstructured":"Tiezheng Yu, Wenliang Dai, Zihan Liu, and Pascale Fung. 2021. Vision guided generative pre-trained language models for multimodal abstractive summarization. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 3995\u20134007."},{"key":"e_1_3_2_59_2","first-page":"1","article-title":"Understanding WeChat user preferences and \u201cWow\u201d diffusion","author":"Zhang Fanjin","year":"2021","unstructured":"Fanjin Zhang, Jie Tang, Xueyi Liu, Zhenyu Hou, Yuxiao Dong, Jing Zhang, Xiao Liu, Ruobing Xie, Kai Zhuang, Xu Zhang, Leyu Lin, and Philip Yu. 2021. Understanding WeChat user preferences and \u201cWow\u201d diffusion. IEEE Trans. Knowl. Data Eng. 34, 12 (2021), 1\u201314.","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"e_1_3_2_60_2","first-page":"695","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Zhang Haoyu","year":"2021","unstructured":"Haoyu Zhang, Meng Liu, Zan Gao, Xiaoqiang Lei, Yinglong Wang, and Liqiang Nie. 2021. Multimodal dialog system: Relational graph-based context-aware question understanding. In Proceedings of the ACM International Conference on Multimedia. ACM, 695\u2013703."},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2020.3031329"},{"key":"e_1_3_2_62_2","first-page":"9604","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Zhang Yichi","year":"2020","unstructured":"Yichi Zhang, Zhijian Ou, and Zhou Yu. 2020. Task-oriented dialog systems that consider multiple appropriate responses under the same context. In Proceedings of the AAAI Conference on Artificial Intelligence. AAAI Press, 9604\u20139611."},{"key":"e_1_3_2_63_2","first-page":"270","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Zhang Yizhe","year":"2020","unstructured":"Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020. DIALOGPT: Large-scale generative pre-training for conversational response generation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 270\u2013278."},{"key":"e_1_3_2_64_2","first-page":"1462","volume-title":"Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Zhu Chenguang","year":"2021","unstructured":"Chenguang Zhu, Ziyi Yang, Robert Gmyr, Michael Zeng, and Xuedong Huang. 2021. Leveraging lead bias for zero-shot abstractive news summarization. In Proceedings of the International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 1462\u20131471."}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3606368","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3606368","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:47:09Z","timestamp":1750178829000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3606368"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,7]]},"references-count":63,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,3,31]]}},"alternative-id":["10.1145\/3606368"],"URL":"https:\/\/doi.org\/10.1145\/3606368","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,7]]},"assertion":[{"value":"2022-08-13","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-14","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}