{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T00:41:59Z","timestamp":1783471319396,"version":"3.55.0"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2024,4,15]],"date-time":"2024-04-15T00:00:00Z","timestamp":1713139200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62222213, U22B2059, 62072423, 62206155"],"award-info":[{"award-number":["62222213, U22B2059, 62072423, 62206155"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"USTC Research Funds of the Double First-Class Initiative","award":["YD2150002009"],"award-info":[{"award-number":["YD2150002009"]}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"crossref","award":["2022M720077"],"award-info":[{"award-number":["2022M720077"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2024,6,30]]},"abstract":"<jats:p>\n            The topic of multimodal conversation systems has recently garnered significant attention across various industries, including travel and retail, among others. While pioneering works in this field have shown promising performance, they often focus solely on context information at the utterance level, overlooking the context-aware dependencies of multimodal semantic elements like words and images. Furthermore, the ordinal information of images, which indicates the relevance between visual context and users\u2019 demands, remains underutilized during the integration of visual content. Additionally, the exploration of how to effectively utilize corresponding attributes provided by users when searching for desired products is still largely unexplored. To address these challenges, we propose PMATE, a\n            <jats:italic>P<\/jats:italic>\n            osition-aware\n            <jats:italic>M<\/jats:italic>\n            ultimodal di\n            <jats:italic>A<\/jats:italic>\n            logue system with seman\n            <jats:italic>T<\/jats:italic>\n            ic\n            <jats:italic>E<\/jats:italic>\n            lements. Specifically, to obtain semantic representations at the element level, we first unfold the multimodal historical utterances and devise a position-aware multimodal element-level encoder. This component considers all images that may be relevant to the current turn and introduces a novel position-aware image selector to choose related images before fusing the information from the two modalities. Finally, we present a knowledge-aware two-stage decoder and an attribute-enhanced image searcher for the tasks of generating textual responses and selecting image responses, respectively. We extensively evaluate our model on two large-scale multimodal dialogue datasets, and the results of our experiments demonstrate that our approach outperforms several baseline methods.\n          <\/jats:p>","DOI":"10.1145\/3645099","type":"journal-article","created":{"date-parts":[[2024,3,12]],"date-time":"2024-03-12T11:48:07Z","timestamp":1710244087000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Multimodal Dialogue Systems via Capturing Context-aware Dependencies and Ordinal Information of Semantic Elements"],"prefix":"10.1145","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1986-9710","authenticated-orcid":false,"given":"Weidong","family":"He","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, Hefei, Anhui, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8061-7486","authenticated-orcid":false,"given":"Zhi","family":"Li","sequence":"additional","affiliation":[{"name":"Tsinghua University, Shenzhen, Guangdong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9921-2078","authenticated-orcid":false,"given":"Hao","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, Anhui, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4246-5386","authenticated-orcid":false,"given":"Tong","family":"Xu","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, Anhui, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6703-2064","authenticated-orcid":false,"given":"Zhefeng","family":"Wang","sequence":"additional","affiliation":[{"name":"HUAWEI Technologies, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9625-2314","authenticated-orcid":false,"given":"Baoxing","family":"Huai","sequence":"additional","affiliation":[{"name":"HUAWEI Technologies, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-3971-4176","authenticated-orcid":false,"given":"Nicholas Jing","family":"Yuan","sequence":"additional","affiliation":[{"name":"HUAWEI Technologies, Hangzhou, Zhejiang, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4835-4102","authenticated-orcid":false,"given":"Enhong","family":"Chen","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, Anhui, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,4,15]]},"reference":[{"key":"e_1_3_2_2_2","article-title":"Layer normalization","author":"Ba Jimmy Lei","year":"2016","unstructured":"Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer normalization. arXiv preprint arXiv:1607.06450 (2016).","journal-title":"arXiv preprint arXiv:1607.06450"},{"key":"e_1_3_2_3_2","doi-asserted-by":"crossref","first-page":"5437","DOI":"10.18653\/v1\/P19-1540","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Chauhan Hardik","year":"2019","unstructured":"Hardik Chauhan, Mauajama Firdaus, Asif Ekbal, and Pushpak Bhattacharyya. 2019. Ordinal and attribute aware response generation in a multimodal dialogue system. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 5437\u20135447."},{"key":"e_1_3_2_4_2","article-title":"Multimodal dialog systems with dual knowledge-enhanced generative pretrained language model","author":"Chen Xiaolin","year":"2022","unstructured":"Xiaolin Chen, Xuemeng Song, Liqiang Jing, Shuo Li, Linmei Hu, and Liqiang Nie. 2022. Multimodal dialog systems with dual knowledge-enhanced generative pretrained language model. arXiv preprint arXiv:2207.07934 (2022).","journal-title":"arXiv preprint arXiv:2207.07934"},{"key":"e_1_3_2_5_2","first-page":"445","volume-title":"Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Cui Chen","year":"2019","unstructured":"Chen Cui, Wenjie Wang, Xuemeng Song, Minlie Huang, Xin-Shun Xu, and Liqiang Nie. 2019. User attention-guided multimodal dialog systems. In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval. 445\u2013454."},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"138","DOI":"10.3115\/1289189.1289273","volume-title":"Proceedings of the 2nd International Conference on Human Language Technology Research","author":"Doddington George","year":"2002","unstructured":"George Doddington. 2002. Automatic evaluation of machine translation quality using n-gram co-occurrence statistics. In Proceedings of the 2nd International Conference on Human Language Technology Research. 138\u2013145."},{"issue":"3","key":"e_1_3_2_7_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3390891","article-title":"Recurrent attention network with reinforced generator for visual dialog","volume":"16","author":"Fan Hehe","year":"2020","unstructured":"Hehe Fan, Linchao Zhu, Yi Yang, and Fei Wu. 2020. Recurrent attention network with reinforced generator for visual dialog. ACM Transactions on Multimedia Computing, Communications, and Applications 16, 3 (2020), 1\u201316.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"issue":"2","key":"e_1_3_2_8_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3430752","article-title":"Aspect-aware response generation for multimodal dialogue system","volume":"12","author":"Firdaus Mauajama","year":"2021","unstructured":"Mauajama Firdaus, Nidhi Thakur, and Asif Ekbal. 2021. Aspect-aware response generation for multimodal dialogue system. ACM Transactions on Intelligent Systems and Technology 12, 2 (2021), 1\u201333.","journal-title":"ACM Transactions on Intelligent Systems and Technology"},{"key":"e_1_3_2_9_2","article-title":"I enjoy writing and playing, do you?: A personalized and emotion grounded dialogue agent using generative adversarial network","author":"Firdaus Mauajama","year":"2022","unstructured":"Mauajama Firdaus, Naveen Thangavelu, Asif Ekbal, and Pushpak Bhattacharyya. 2022. I enjoy writing and playing, do you?: A personalized and emotion grounded dialogue agent using generative adversarial network. IEEE Transactions on Affective Computing. Published Online, February 28, 2022.","journal-title":"IEEE Transactions on Affective Computing."},{"issue":"5","key":"e_1_3_2_10_2","doi-asserted-by":"crossref","first-page":"378","DOI":"10.1037\/h0031619","article-title":"Measuring nominal scale agreement among many raters.","volume":"76","author":"Fleiss Joseph L.","year":"1971","unstructured":"Joseph L. Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychological Bulletin 76, 5 (1971), 378.","journal-title":"Psychological Bulletin"},{"issue":"2","key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3429980","article-title":"Learning to respond with your favorite stickers: A framework of unifying multi-modality and user preference in multi-turn dialog","volume":"39","author":"Gao Shen","year":"2021","unstructured":"Shen Gao, Xiuying Chen, Li Liu, Dongyan Zhao, and Rui Yan. 2021. Learning to respond with your favorite stickers: A framework of unifying multi-modality and user preference in multi-turn dialog. ACM Transactions on Information Systems 39, 2 (2021), 1\u201332.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_2_12_2","first-page":"770","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 770\u2013778."},{"key":"e_1_3_2_13_2","first-page":"2755","volume-title":"Proceedings of the 28th ACM International Conference on Multimedia","author":"He Weidong","year":"2020","unstructured":"Weidong He, Zhi Li, Dongcai Lu, Enhong Chen, Tong Xu, Baoxing Huai, and Jing Yuan. 2020. Multimodal dialogue systems via capturing context-aware dependencies of semantic elements. In Proceedings of the 28th ACM International Conference on Multimedia. 2755\u20132764."},{"issue":"11","key":"e_1_3_2_14_2","doi-asserted-by":"crossref","first-page":"2072","DOI":"10.1109\/TASLP.2018.2852492","article-title":"Cross-language neural dialog state tracker for large ontologies using hierarchical attention","volume":"26","author":"Jang Youngsoo","year":"2018","unstructured":"Youngsoo Jang, Jiyeon Ham, Byung-Jun Lee, and Kee-Eung Kim. 2018. Cross-language neural dialog state tracker for large ontologies using hierarchical attention. IEEE\/ACM Transactions on Audio, Speech, and Language Processing 26, 11 (2018), 2072\u20132082.","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"issue":"5","key":"e_1_3_2_15_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3453154","article-title":"A survey on conversational recommender systems","volume":"54","author":"Jannach Dietmar","year":"2021","unstructured":"Dietmar Jannach, Ahtsham Manzoor, Wanling Cai, and Li Chen. 2021. A survey on conversational recommender systems. ACM Computing Surveys 54, 5 (2021), 1\u201336.","journal-title":"ACM Computing Surveys"},{"key":"e_1_3_2_16_2","article-title":"An information retrieval approach to short text conversation","author":"Ji Zongcheng","year":"2014","unstructured":"Zongcheng Ji, Zhengdong Lu, and Hang Li. 2014. An information retrieval approach to short text conversation. arXiv preprint arXiv:1408.6988 (2014).","journal-title":"arXiv preprint arXiv:1408.6988"},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","first-page":"2437","DOI":"10.1109\/TASLP.2021.3077119","article-title":"Randomly wired network based on RoBERTa and dialog history attention for response selection","volume":"29","author":"Kim Byoungjae","year":"2021","unstructured":"Byoungjae Kim, Jungyun Seo, and Myoung-Wan Koo. 2021. Randomly wired network based on RoBERTa and dialog history attention for response selection. IEEE\/ACM Transactions on Audio, Speech, and Language Processing 29 (2021), 2437\u20132442.","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"e_1_3_2_18_2","first-page":"304","volume-title":"Proceedings of the 13th International Conference on Web Search and Data Mining","author":"Lei Wenqiang","year":"2020","unstructured":"Wenqiang Lei, Xiangnan He, Yisong Miao, Qingyun Wu, Richang Hong, Min-Yen Kan, and Tat-Seng Chua. 2020. Estimation-action-reflection: Towards deep interaction between conversational and recommender systems. In Proceedings of the 13th International Conference on Web Search and Data Mining. 304\u2013312."},{"key":"e_1_3_2_19_2","first-page":"1437","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Lei Wenqiang","year":"2018","unstructured":"Wenqiang Lei, Xisen Jin, Min-Yen Kan, Zhaochun Ren, Xiangnan He, and Dawei Yin. 2018. Sequicity: Simplifying task-oriented dialogue systems with single sequence-to-sequence architectures. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1437\u20131447."},{"issue":"1","key":"e_1_3_2_20_2","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1109\/89.817450","article-title":"A stochastic model of human-machine interaction for learning dialog strategies","volume":"8","author":"Levin Esther","year":"2000","unstructured":"Esther Levin, Roberto Pieraccini, and Wieland Eckert. 2000. A stochastic model of human-machine interaction for learning dialog strategies. IEEE Transactions on Speech and Audio Processing 8, 1 (2000), 11\u201323.","journal-title":"IEEE Transactions on Speech and Audio Processing"},{"issue":"1","key":"e_1_3_2_21_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1126004.1126005","article-title":"Content-based multimedia information retrieval: State of the art and challenges","volume":"2","author":"Lew Michael S.","year":"2006","unstructured":"Michael S. Lew, Nicu Sebe, Chabane Djeraba, and Ramesh Jain. 2006. Content-based multimedia information retrieval: State of the art and challenges. ACM Transactions on Multimedia Computing, Communications, and Applications 2, 1 (2006), 1\u201319.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_2_22_2","first-page":"2157","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201917)","author":"Li Jiwei","year":"2017","unstructured":"Jiwei Li, Will Monroe, Tianlin Shi, S\u00e9bastien Jean, Alan Ritter, and Dan Jurafsky. 2017. Adversarial learning for neural dialogue generation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201917). 2157\u20132169."},{"issue":"4","key":"e_1_3_2_23_2","first-page":"1","article-title":"Densely enhanced semantic network for conversation system in social media","volume":"18","author":"Li Yongrui","year":"2022","unstructured":"Yongrui Li, Zengfu Wang, and Jun Yu. 2022. Densely enhanced semantic network for conversation system in social media. ACM Transactions on Multimedia Computing, Communications, and Applications 18, 4 (2022), 1\u201324.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_2_24_2","doi-asserted-by":"crossref","first-page":"675","DOI":"10.1145\/3404835.3462970","volume-title":"Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Liao Lizi","year":"2021","unstructured":"Lizi Liao, Le Hong Long, Zheng Zhang, Minlie Huang, and Tat-Seng Chua. 2021. MMConv: An environment for multimodal conversational search across multiple domains. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 675\u2013684."},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","first-page":"801","DOI":"10.1145\/3240508.3240605","volume-title":"Proceedings of the 26th ACM International Conference on Multimedia","author":"Liao Lizi","year":"2018","unstructured":"Lizi Liao, Yunshan Ma, Xiangnan He, Richang Hong, and Tat-Seng Chua. 2018. Knowledge-aware multimodal dialogue systems. In Proceedings of the 26th ACM International Conference on Multimedia. 801\u2013809."},{"issue":"5","key":"e_1_3_2_26_2","doi-asserted-by":"crossref","first-page":"2485","DOI":"10.1109\/TKDE.2020.3008563","article-title":"Topic-guided conversational recommender in multiple domains","volume":"34","author":"Liao Lizi","year":"2022","unstructured":"Lizi Liao, Ryuichi Takanobu, Yunshan Ma, Xun Yang, Minlie Huang, and Tat-Seng Chua. 2022. Topic-guided conversational recommender in multiple domains. IEEE Transactions on Knowledge and Data Engineering 34, 5 (2022), 2485\u20132496.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"issue":"4","key":"e_1_3_2_27_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3498340","article-title":"Answer questions with right image regions: A visual attention regularization approach","volume":"18","author":"Liu Yibing","year":"2022","unstructured":"Yibing Liu, Yangyang Guo, Jianhua Yin, Xuemeng Song, Weifeng Liu, Liqiang Nie, and Min Zhang. 2022. Answer questions with right image regions: A visual attention regularization approach. ACM Transactions on Multimedia Computing, Communications, and Applications 18, 4 (2022), 1\u201318.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_2_28_2","article-title":"Graph-grounded goal planning for conversational recommendation","author":"Liu Zeming","year":"2022","unstructured":"Zeming Liu, Ding Zhou, Hao Liu, Haifeng Wang, Zheng-Yu Niu, Hua Wu, Wanxiang Che, Ting Liu, and Hui Xiong. 2022. Graph-grounded goal planning for conversational recommendation. IEEE Transactions on Knowledge and Data Engineering. Published Online, February 1, 2022.","journal-title":"IEEE Transactions on Knowledge and Data Engineering."},{"key":"e_1_3_2_29_2","first-page":"103","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Ma Zhiyuan","year":"2022","unstructured":"Zhiyuan Ma, Jianjun Li, Guohui Li, and Yongjing Cheng. 2022. UniTranSeR: A unified transformer semantic representation framework for multimodal task-oriented dialog system. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 103\u2013114."},{"issue":"3","key":"e_1_3_2_30_2","doi-asserted-by":"crossref","first-page":"530","DOI":"10.1109\/TASLP.2014.2383614","article-title":"Using recurrent neural networks for slot filling in spoken language understanding","volume":"23","author":"Mesnil Gr\u00e9goire","year":"2014","unstructured":"Gr\u00e9goire Mesnil, Yann Dauphin, Kaisheng Yao, Yoshua Bengio, Li Deng, Dilek Hakkani-Tur, Xiaodong He, Larry Heck, Gokhan Tur, Dong Yu, and Geoffrey Zweig. 2014. Using recurrent neural networks for slot filling in spoken language understanding. IEEE\/ACM Transactions on Audio, Speech, and Language Processing 23, 3 (2014), 530\u2013539.","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"e_1_3_2_31_2","doi-asserted-by":"crossref","first-page":"7732","DOI":"10.1109\/TIP.2021.3108724","article-title":"Conversational image search","volume":"30","author":"Nie Liqiang","year":"2021","unstructured":"Liqiang Nie, Fangkai Jiao, Wenjie Wang, Yinglong Wang, and Qi Tian. 2021. Conversational image search. IEEE Transactions on Image Processing 30 (2021), 7732\u20137743.","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_3_2_32_2","first-page":"1098","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Nie Liqiang","year":"2019","unstructured":"Liqiang Nie, Wenjie Wang, Richang Hong, Meng Wang, and Qi Tian. 2019. Multimodal dialog system: Generating responses via adaptive decoders. In Proceedings of the ACM International Conference on Multimedia. 1098\u20131106."},{"key":"e_1_3_2_33_2","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","first-page":"2182","DOI":"10.18653\/v1\/P18-1203","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Peng Baolin","year":"2018","unstructured":"Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Kam-Fai Wong. 2018. Deep Dyna-Q: Integrating planning for task-completion dialogue policy learning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2182\u20132192."},{"key":"e_1_3_2_35_2","first-page":"2231","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing","author":"Peng Baolin","year":"2017","unstructured":"Baolin Peng, Xiujun Li, Lihong Li, Jianfeng Gao, Asli Celikyilmaz, Sungjin Lee, and Kam-Fai Wong. 2017. Composite task-completion dialogue policy learning via hierarchical deep reinforcement learning. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2231\u20132240."},{"key":"e_1_3_2_36_2","article-title":"Few-shot natural language generation for task-oriented dialog","author":"Peng Baolin","year":"2020","unstructured":"Baolin Peng, Chenguang Zhu, Chunyuan Li, Xiujun Li, Jinchao Li, Michael Zeng, and Jianfeng Gao. 2020. Few-shot natural language generation for task-oriented dialog. arXiv preprint arXiv:2002.12328 (2020).","journal-title":"arXiv preprint arXiv:2002.12328"},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","first-page":"1532","DOI":"10.3115\/v1\/D14-1162","volume-title":"Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201914)","author":"Pennington Jeffrey","year":"2014","unstructured":"Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP\u201914). 1532\u20131543."},{"key":"e_1_3_2_38_2","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence, the 30th Innovative Applications of Arti","author":"Saha Amrita","year":"2018","unstructured":"Amrita Saha, Mitesh M. Khapra, and Karthik Sankaranarayanan. 2018. Towards building large scale multimodal domain-aware conversation systems. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence, the 30th Innovative Applications of Artificial Intelligence Conference, and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (AAAI\/IAAI\/EAAI\u201918). 696\u2013704."},{"key":"e_1_3_2_39_2","article-title":"Hierarchical neural network generative models for movie dialogues","author":"Serban Iulian V.","year":"2015","unstructured":"Iulian V. Serban, Alessandro Sordoni, Yoshua Bengio, Aaron Courville, and Joelle Pineau. 2015. Hierarchical neural network generative models for movie dialogues. arXiv preprint arXiv:1507.04808 (2015).","journal-title":"arXiv preprint arXiv:1507.04808"},{"key":"e_1_3_2_40_2","volume-title":"Proceedings of the 30th AAAI Conference on Artificial Intelligence (AAAI\u201916)","author":"Serban Iulian V.","year":"2016","unstructured":"Iulian V. Serban, Alessandro Sordoni, Yoshua Bengio, Aaron Courville, and Joelle Pineau. 2016. Building end-to-end dialogue systems using generative hierarchical neural network models. In Proceedings of the 30th AAAI Conference on Artificial Intelligence (AAAI\u201916). 3776\u20133783."},{"key":"e_1_3_2_41_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"32","author":"Shen Xiaoyu","year":"2018","unstructured":"Xiaoyu Shen, Hui Su, Shuzi Niu, and Vera Demberg. 2018. Improving variational encoder-decoders in dialogue generation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 32."},{"key":"e_1_3_2_42_2","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014).","journal-title":"arXiv preprint arXiv:1409.1556"},{"key":"e_1_3_2_43_2","first-page":"6582","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201922)","author":"Sun Rongyi","year":"2022","unstructured":"Rongyi Sun, Borun Chen, Qingyu Zhou, Yinghui Li, Yunbo Cao, and Hai-Tao Zheng. 2022. A non-hierarchical attention network with modality dropout for textual response generation in multimodal dialogue systems. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP\u201922). IEEE, 6582\u20136586."},{"key":"e_1_3_2_44_2","first-page":"235","volume-title":"Proceedings of the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Sun Yueming","year":"2018","unstructured":"Yueming Sun and Yi Zhang. 2018. Conversational recommender system. In Proceedings of the 41st International ACM SIGIR Conference on Research and Development in Information Retrieval. 235\u2013244."},{"key":"e_1_3_2_45_2","article-title":"Sequence to sequence learning with neural networks","volume":"27","author":"Sutskever Ilya","year":"2014","unstructured":"Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. Advances in Neural Information Processing Systems 27 (2014), 1\u20139.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_46_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing systems 30 (2017), 1\u201311.","journal-title":"Advances in Neural Information Processing systems"},{"key":"e_1_3_2_47_2","article-title":"A neural conversational model","author":"Vinyals Oriol","year":"2015","unstructured":"Oriol Vinyals and Quoc Le. 2015. A neural conversational model. arXiv preprint arXiv:1506.05869 (2015).","journal-title":"arXiv preprint arXiv:1506.05869"},{"issue":"2","key":"e_1_3_2_48_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3463913","article-title":"HyperSoRec: Exploiting hyperbolic user and item representations with multiple aspects for social-aware recommendation","volume":"40","author":"Wang Hao","year":"2021","unstructured":"Hao Wang, Defu Lian, Hanghang Tong, Qi Liu, Zhenya Huang, and Enhong Chen. 2021. HyperSoRec: Exploiting hyperbolic user and item representations with multiple aspects for social-aware recommendation. ACM Transactions on Information Systems 40, 2 (2021), 1\u201328.","journal-title":"ACM Transactions on Information Systems"},{"key":"e_1_3_2_49_2","first-page":"1711","volume-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing","author":"Wen Tsung-Hsien","year":"2015","unstructured":"Tsung-Hsien Wen, Milica Gasic, Nikola Mrk\u0161i\u0107, Pei-Hao Su, David Vandyke, and Steve Young. 2015. Semantically conditioned LSTM-based natural language generation for spoken dialogue systems. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 1711\u20131721."},{"key":"e_1_3_2_50_2","first-page":"438","volume-title":"Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wen Tsung-Hsien","year":"2017","unstructured":"Tsung-Hsien Wen, David Vandyke, Nikola Mrk\u0161i\u0107, Milica Gasic, Lina M. Rojas Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young. 2017. A network-based end-to-end trainable task-oriented dialogue system. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 438\u2013449."},{"issue":"2","key":"e_1_3_2_51_2","doi-asserted-by":"crossref","first-page":"270","DOI":"10.1162\/neco.1989.1.2.270","article-title":"A learning algorithm for continually running fully recurrent neural networks","volume":"1","author":"Williams Ronald J.","year":"1989","unstructured":"Ronald J. Williams and David Zipser. 1989. A learning algorithm for continually running fully recurrent neural networks. Neural Computation 1, 2 (1989), 270\u2013280.","journal-title":"Neural Computation"},{"key":"e_1_3_2_52_2","article-title":"A survey on large language models for recommendation","author":"Wu Likang","year":"2023","unstructured":"Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, Lui Xiong, and Enhong Chen. 2023. A survey on large language models for recommendation. arXiv preprint arXiv:2305.19860 (2023).","journal-title":"arXiv preprint arXiv:2305.19860"},{"key":"e_1_3_2_53_2","article-title":"Deliberation networks: Sequence generation beyond one-pass decoding","volume":"30","author":"Xia Yingce","year":"2017","unstructured":"Yingce Xia, Fei Tian, Lijun Wu, Jianxin Lin, Tao Qin, Nenghai Yu, and Tie-Yan Liu. 2017. Deliberation networks: Sequence generation beyond one-pass decoding. Advances in Neural Information Processing Systems 30 (2017), 1\u201311.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_54_2","first-page":"7338","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Xiong Hao","year":"2019","unstructured":"Hao Xiong, Zhongjun He, Hua Wu, and Haifeng Wang. 2019. Modeling coherence for discourse neural machine translation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 7338\u20137345."},{"key":"e_1_3_2_55_2","doi-asserted-by":"crossref","first-page":"839","DOI":"10.1145\/3404835.3462881","volume-title":"Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Yuan Yifei","year":"2021","unstructured":"Yifei Yuan and Wai Lam. 2021. Conversational fashion image retrieval via multiturn natural language feedback. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 839\u2013848."},{"key":"e_1_3_2_56_2","doi-asserted-by":"crossref","first-page":"2578","DOI":"10.1145\/3442381.3450026","volume-title":"Proceedings of the Web Conference 2021","author":"Zeng Zengfeng","year":"2021","unstructured":"Zengfeng Zeng, Dan Ma, Haiqin Yang, Zhen Gou, and Jianping Shen. 2021. Automatic intent-slot induction for dialogue systems. In Proceedings of the Web Conference 2021. 2578\u20132589."},{"key":"e_1_3_2_57_2","doi-asserted-by":"crossref","first-page":"695","DOI":"10.1145\/3474085.3475234","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia","author":"Zhang Haoyu","year":"2021","unstructured":"Haoyu Zhang, Meng Liu, Zan Gao, Xiaoqiang Lei, Yinglong Wang, and Liqiang Nie. 2021. Multimodal dialog system: Relational graph-based context-aware question understanding. In Proceedings of the 29th ACM International Conference on Multimedia. 695\u2013703."},{"key":"e_1_3_2_58_2","doi-asserted-by":"crossref","first-page":"177","DOI":"10.1145\/3269206.3271776","volume-title":"Proceedings of the 27th ACM International Conference on Information and Knowledge Management","author":"Zhang Yongfeng","year":"2018","unstructured":"Yongfeng Zhang, Xu Chen, Qingyao Ai, Liu Yang, and W. Bruce Croft. 2018. Towards conversational search and recommendation: System ask, user respond. In Proceedings of the 27th ACM International Conference on Information and Knowledge Management. 177\u2013186."},{"issue":"10","key":"e_1_3_2_59_2","doi-asserted-by":"crossref","first-page":"2011","DOI":"10.1007\/s11431-020-1692-3","article-title":"Recent advances and challenges in task-oriented dialog systems","volume":"63","author":"Zhang Zheng","year":"2020","unstructured":"Zheng Zhang, Ryuichi Takanobu, Qi Zhu, MinLie Huang, and XiaoYan Zhu. 2020. Recent advances and challenges in task-oriented dialog systems. Science China Technological Sciences 63, 10 (2020), 2011\u20132027.","journal-title":"Science China Technological Sciences"},{"key":"e_1_3_2_60_2","doi-asserted-by":"crossref","first-page":"654","DOI":"10.18653\/v1\/P17-1061","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhao Tiancheng","year":"2017","unstructured":"Tiancheng Zhao, Ran Zhao, and Maxine Eskenazi. 2017. Learning discourse-level diversity for neural dialog models using conditional variational autoencoders. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 654\u2013664."},{"key":"e_1_3_2_61_2","doi-asserted-by":"crossref","first-page":"1006","DOI":"10.1145\/3394486.3403143","volume-title":"Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Zhou Kun","year":"2020","unstructured":"Kun Zhou, Wayne Xin Zhao, Shuqing Bian, Yuanhang Zhou, Ji-Rong Wen, and Jingsong Yu. 2020. Improving conversational recommender systems via knowledge graph based semantic fusion. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 1006\u20131014."}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3645099","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3645099","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:03:27Z","timestamp":1750291407000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3645099"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,15]]},"references-count":60,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,6,30]]}},"alternative-id":["10.1145\/3645099"],"URL":"https:\/\/doi.org\/10.1145\/3645099","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,15]]},"assertion":[{"value":"2022-11-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-01-16","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}