{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,11]],"date-time":"2025-12-11T03:05:21Z","timestamp":1765422321818,"version":"3.41.0"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,5,31]],"date-time":"2023-05-31T00:00:00Z","timestamp":1685491200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Ministry of Education INDIA","award":["1-3146198040"],"award-info":[{"award-number":["1-3146198040"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:p>In this article, we have defined a novel task of affective feedback synthesis that generates feedback for input text and corresponding images in a way similar to humans responding to multimodal data. A feedback synthesis system has been proposed and trained using ground-truth human comments along with image\u2013text input. We have also constructed a large-scale dataset consisting of images, text, Twitter user comments, and the number of likes for the comments by crawling news articles through Twitter feeds. The proposed system extracts textual features using a transformer-based textual encoder. The visual features have been extracted using a Faster region-based convolutional neural networks model. The textual and visual features have been concatenated to construct multimodal features that the decoder uses to synthesize the feedback. We have compared the results of the proposed system with baseline models using quantitative and qualitative measures. The synthesized feedbacks have been analyzed using automatic and human evaluation. They have been found to be semantically similar to the ground-truth comments and relevant to the given text\u2013image input.<\/jats:p>","DOI":"10.1145\/3589186","type":"journal-article","created":{"date-parts":[[2023,3,24]],"date-time":"2023-03-24T12:14:52Z","timestamp":1679660092000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Affective Feedback Synthesis Towards Multimodal Text and Image Data"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4318-1353","authenticated-orcid":false,"given":"Puneet","family":"Kumar","sequence":"first","affiliation":[{"name":"Indian Institute of Technology Roorkee, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0980-8765","authenticated-orcid":false,"given":"Gaurav","family":"Bhatt","sequence":"additional","affiliation":[{"name":"Indian Institute of Technology Hyderabad, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6310-880X","authenticated-orcid":false,"given":"Omkar","family":"Ingle","sequence":"additional","affiliation":[{"name":"Indian Institute of Technology Roorkee, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8999-9142","authenticated-orcid":false,"given":"Daksh","family":"Goyal","sequence":"additional","affiliation":[{"name":"National Institute of Technology Karnataka, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6277-6267","authenticated-orcid":false,"given":"Balasubramanian","family":"Raman","sequence":"additional","affiliation":[{"name":"Indian Institute of Technology Roorkee, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,5,31]]},"reference":[{"key":"e_1_3_3_2_2","first-page":"7558","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201919)","author":"Alamri Huda","year":"2019","unstructured":"Huda Alamri, Vincent Cartillier, Abhishek Das, Jue Wang, Anoop Cherian, Irfan Essa, Dhruv Batra, Tim K. Marks, Chiori Hori, Peter Anderson, et\u00a0al. 2019. Audio visual scene aware dialog. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201919). 7558\u20137567."},{"key":"e_1_3_3_3_2","volume-title":"Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14","author":"Anderson Peter","year":"2016","unstructured":"Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016. SPICE: Semantic propositional image caption evaluation. Computer Vision\u2013ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14 , Springer, 382\u2013398."},{"key":"e_1_3_3_4_2","first-page":"9347","volume-title":"Conference on Empirical Methods in Natural Language Processing (EMNLP\u201920)","author":"Bhandari Manik","year":"2020","unstructured":"Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020. Re-evaluating evaluation in text summarization. In Conference on Empirical Methods in Natural Language Processing (EMNLP\u201920). 9347\u20139359."},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.18608\/jla.2016.32.11"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.tele.2013.06.003"},{"key":"e_1_3_3_7_2","first-page":"7504","volume-title":"34th AAAI Conference on Artificial Intelligence (AAAI)","volume":"34","author":"Chen Feilong","year":"2020","unstructured":"Feilong Chen, Fandong Meng, Jiaming Xu, Peng Li, Bo Xu, and Jie Zhou. 2020. DMRM: A dual channel multi hop reasoning model for visual dialog. In 34th AAAI Conference on Artificial Intelligence (AAAI), Vol. 34. 7504\u20137511."},{"key":"e_1_3_3_8_2","doi-asserted-by":"crossref","first-page":"4046","DOI":"10.18653\/v1\/D18-1438","volume-title":"Conference on Empirical Methods in Natural Language Processing (EMNLP\u201918)","author":"Chen Jingqiang","year":"2018","unstructured":"Jingqiang Chen and Hai Zhuge. 2018. Abstractive text-image summarization using multimodal attention hierarchical RNN. In Conference on Empirical Methods in Natural Language Processing (EMNLP\u201918). 4046\u20134056."},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2015.2432578"},{"key":"e_1_3_3_10_2","first-page":"1931","volume-title":"International Conference on Machine Learning (ICML\u201921)","author":"Cho Jaemin","year":"2021","unstructured":"Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. 2021. Unifying vision-and-language tasks via text generation. In International Conference on Machine Learning (ICML\u201921). 1931\u20131942."},{"key":"e_1_3_3_11_2","doi-asserted-by":"crossref","DOI":"10.1007\/978-0-387-39940-9_488","article-title":"Mean reciprocal rank","volume":"1703","author":"Craswell Nick","year":"2009","unstructured":"Nick Craswell. 2009. Mean reciprocal rank. Encyclopedia of Database Systems 1703, Springer. https:\/\/www.bibsonomy.org\/bibtex\/2e8204c7513a94917d206b34004d87a54\/dblp.","journal-title":"Encyclopedia of Database Systems"},{"key":"e_1_3_3_12_2","unstructured":"Benjamin Dornel. 2020. New York Times Articles & Comments Dataset. Retrieved February 20 2022 from https:\/\/www.kaggle.com\/benjaminawd\/new-york-times-articles-comments-2020. (2020)."},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2020.113679"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2019.10.044"},{"key":"e_1_3_3_15_2","volume-title":"29th International Joint Conference on Artificial Intelligence (IJCAI\u201920)","author":"Gao Shen","year":"2020","unstructured":"Shen Gao, Xiuying Chen, Zhaochun Ren, Dongyan Zhao, and Rui Yan. 2020. From standard summarization to new tasks and beyond: Summarization with manifold information. In 29th International Joint Conference on Artificial Intelligence (IJCAI\u201920)."},{"key":"e_1_3_3_16_2","first-page":"1440","volume-title":"IEEE\/CVF International Conference on Computer Vision (ICCV\u201915)","author":"Girshick Ross","year":"2015","unstructured":"Ross Girshick. 2015. Fast R-CNN. In IEEE\/CVF International Conference on Computer Vision (ICCV\u201915). 1440\u20131448."},{"key":"e_1_3_3_17_2","doi-asserted-by":"crossref","first-page":"580","DOI":"10.1109\/CVPR.2014.81","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201914)","author":"Girshick Ross","year":"2014","unstructured":"Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. 2014. Rich feature hierarchies for accurate object detection and semantic segmentation. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201914). 580\u2013587."},{"key":"e_1_3_3_18_2","volume-title":"International Conference on Learning Representations (ICLR\u201919)","author":"Gu Xiaodong","year":"2019","unstructured":"Xiaodong Gu, Kyunghyun Cho, Jung-Woo Ha, and Sunghun Kim. 2019. DialogWAE: Multimodal response generation with conditional Wasserstein auto encoder. In International Conference on Learning Representations (ICLR\u201919)."},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2017.2741510"},{"issue":"10","key":"e_1_3_3_20_2","first-page":"2487726","article-title":"Using innovative instructions to create trustworthy software solutions.","volume":"11","author":"Hoekstra Matthew","year":"2013","unstructured":"Matthew Hoekstra, Reshma Lal, Pradeep Pappachan, Vinay Phegade, and Juan Del Cuvillo. 2013. Using innovative instructions to create trustworthy software solutions. Hardware and Architectural Support for Security and Privacy, ISCA 11, 10.1145 (2013), 2487726\u20132488370.","journal-title":"Hardware and Architectural Support for Security and Privacy, ISCA"},{"key":"e_1_3_3_21_2","first-page":"2352","volume-title":"44th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201919)","author":"Hori Chiori","year":"2019","unstructured":"Chiori Hori, Huda Alamri, Jue Wang, Gordon Wichern, Takaaki Hori, Anoop Cherian, Tim K. Marks, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das, et\u00a0al. 2019. End-to-end audio visual scene aware dialog using multimodal attention-based video features. In 44th IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP\u201919). 2352\u20132356."},{"key":"e_1_3_3_22_2","first-page":"350","volume-title":"ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD\u201918)","author":"Hu Anthony","year":"2018","unstructured":"Anthony Hu and Seth Flaxman. 2018. Multimodal sentiment analysis to explore the structure of emotions. In ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD\u201918). 350\u2013358."},{"key":"e_1_3_3_23_2","first-page":"7969","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Hu Junjie","year":"2020","unstructured":"Junjie Hu, Yu Cheng, Zhe Gan, Jingjing Liu, Jianfeng Gao, and Graham Neubig. 2020. What makes a good story? Designing composite rewards for visual storytelling. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 7969\u20137976."},{"key":"e_1_3_3_24_2","first-page":"1233","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL\u201916)","author":"Huang Ting-Hao","year":"2016","unstructured":"Ting-Hao Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, et\u00a0al. 2016. Visual storytelling. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics (NAACL\u201916). 1233\u20131239."},{"key":"e_1_3_3_25_2","first-page":"11125","volume-title":"34th AAAI Conference on Artificial Intelligence (AAAI\u201920)","volume":"34","author":"Jiang Xiaoze","year":"2020","unstructured":"Xiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang, Xingxing Zhang, Yue Hu, and Qi Wu. 2020. DualVD: An adaptive dual encoding model for deep visual understanding in visual dialogue. In 34th AAAI Conference on Artificial Intelligence (AAAI\u201920), Vol. 34. 11125\u201311132."},{"key":"e_1_3_3_26_2","first-page":"2024","volume-title":"Conference on Empirical Methods in Natural Language Processing (EMNLP\u201919)","author":"Kang Gi-Cheon","year":"2019","unstructured":"Gi-Cheon Kang, Jaeseo Lim, and Byoung-Tak Zhang. 2019. Dual attention networks for visual reference resolution in visual dialog. In Conference on Empirical Methods in Natural Language Processing (EMNLP\u201919). 2024\u20132033."},{"key":"e_1_3_3_27_2","first-page":"9332","volume-title":"Conference on Empirical Methods in Natural Language Processing (EMNLP\u201920)","author":"Kry\u015bci\u0144ski Wojciech","year":"2020","unstructured":"Wojciech Kry\u015bci\u0144ski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the factual consistency of abstractive text summarization. In Conference on Empirical Methods in Natural Language Processing (EMNLP\u201920). 9332\u20139346."},{"key":"e_1_3_3_28_2","first-page":"314","volume-title":"International Conference on Image Processing (ICIP\u201921)","author":"Kumar Puneet","year":"2021","unstructured":"Puneet Kumar, Vedanti Khokher, Yukti Gupta, and Balasubramanian Raman. 2021. Hybrid fusion based approach for multimodal emotion recognition with insufficient labeled data. In International Conference on Image Processing (ICIP\u201921). IEEE, 314\u2013318."},{"key":"e_1_3_3_29_2","article-title":"WikiLingua: A new benchmark dataset for cross-lingual abstractive summarization","author":"Ladhak Faisal","year":"2020","unstructured":"Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown. 2020. WikiLingua: A new benchmark dataset for cross-lingual abstractive summarization. arXiv preprint arXiv:2010.03093 (2020).","journal-title":"arXiv preprint arXiv:2010.03093"},{"key":"e_1_3_3_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2015.2460697"},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10590-009-9059-4"},{"key":"e_1_3_3_32_2","first-page":"8188","volume-title":"34th AAAI Conference on Artificial Intelligence (AAAI\u201920)","volume":"34","author":"Li Haoran","year":"2020","unstructured":"Haoran Li, Peng Yuan, Song Xu, Youzheng Wu, Xiaodong He, and Bowen Zhou. 2020. Aspect aware multimodal summarization for Chinese e-commerce products. In 34th AAAI Conference on Artificial Intelligence (AAAI\u201920), Vol. 34. 8188\u20138195."},{"key":"e_1_3_3_33_2","first-page":"297","volume-title":"Proceedings of the International Conference on Multimedia Retrieval (ICMR\u201919)","author":"Li Nanxing","year":"2019","unstructured":"Nanxing Li, Bei Liu, Zhizhong Han, Yu-Shen Liu, and Jianlong Fu. 2019. Emotion reinforced visual storytelling. In Proceedings of the International Conference on Multimedia Retrieval (ICMR\u201919). 297\u2013305."},{"issue":"3","key":"e_1_3_3_34_2","first-page":"33","article-title":"Multimedia news summarization in search","volume":"7","author":"Li Zechao","year":"2016","unstructured":"Zechao Li, Jinhui Tang, Xueming Wang, Jing Liu, and Hanqing Lu. 2016. Multimedia news summarization in search. ACM Transactions on Intelligent Systems and Technology (TIST\u201916) 7, 3 (2016), 33.","journal-title":"ACM Transactions on Intelligent Systems and Technology (TIST\u201916)"},{"key":"e_1_3_3_35_2","first-page":"74","volume-title":"Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out. 74\u201381."},{"key":"e_1_3_3_36_2","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692 (2019).","journal-title":"arXiv preprint arXiv:1907.11692"},{"key":"e_1_3_3_37_2","volume-title":"IEEE Automatic Speech Recognition and Understanding Workshop (ASRU\u201921)","author":"Makiuchi Mariana Rodrigues","year":"2021","unstructured":"Mariana Rodrigues Makiuchi, Kuniaki Uto, and Koichi Shinoda. 2021. Multimodal emotion recognition with high-level speech and text features. In IEEE Automatic Speech Recognition and Understanding Workshop (ASRU\u201921)."},{"issue":"3","key":"e_1_3_3_38_2","doi-asserted-by":"crossref","first-page":"288","DOI":"10.1007\/s40596-021-01425-y","article-title":"Increasing resident support following patient suicide: Assessing resident perceptions of a longitudinal, multimodal patient suicide curriculum","volume":"45","author":"McCutcheon Samar","year":"2021","unstructured":"Samar McCutcheon and Julie Hyman. 2021. Increasing resident support following patient suicide: Assessing resident perceptions of a longitudinal, multimodal patient suicide curriculum. Academic Psychiatry 45, 3 (2021), 288\u2013291.","journal-title":"Academic Psychiatry"},{"key":"e_1_3_3_39_2","doi-asserted-by":"crossref","unstructured":"Gary McDarby James Condron Darran Hughes and Ned Augenblick. 2003. Affective feedback. Enabling Technologies (2003).","DOI":"10.1016\/B978-0-443-07247-5.50010-9"},{"issue":"1","key":"e_1_3_3_40_2","doi-asserted-by":"crossref","first-page":"36","DOI":"10.1109\/TAFFC.2019.2902091","article-title":"Recognizing induced emotions of movie audiences from multimodal information","volume":"12","author":"Muszynski Michal","year":"2019","unstructured":"Michal Muszynski, Leimin Tian, Catherine Lai, Johanna D. Moore, Theodoros Kostoulas, Patrizia Lombardo, Thierry Pun, and Guillaume Chanel. 2019. Recognizing induced emotions of movie audiences from multimodal information. IEEE Transactions on Affective Computing 12, 1 (2019), 36\u201352.","journal-title":"IEEE Transactions on Affective Computing"},{"key":"e_1_3_3_41_2","article-title":"Neural extractive summarization with side information","author":"Narayan Shashi","year":"2017","unstructured":"Shashi Narayan, Nikos Papasarantopoulos, Shay B. Cohen, and Mirella Lapata. 2017. Neural extractive summarization with side information. arXiv preprint arXiv:1704.04530 (2017).","journal-title":"arXiv preprint arXiv:1704.04530"},{"key":"e_1_3_3_42_2","first-page":"6679","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201919)","author":"Niu Yulei","year":"2019","unstructured":"Yulei Niu, Hanwang Zhang, Manli Zhang, Jianhong Zhang, Zhiwu Lu, and Ji-Rong Wen. 2019. Recursive visual attention in visual dialog. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201919). 6679\u20136688."},{"key":"e_1_3_3_43_2","first-page":"3","volume-title":"ACM International Conference on Multimedia Retrieval (ICLR\u201919)","author":"Fortin Mathieu Page","year":"2019","unstructured":"Mathieu Page Fortin and Brahim Chaib-draa. 2019. Multimodal multitask emotion recognition using images, texts and tags. In ACM International Conference on Multimedia Retrieval (ICLR\u201919). 3\u201310."},{"key":"e_1_3_3_44_2","article-title":"Content-based visual summarization for image collections","author":"Pan Xingjia","year":"2019","unstructured":"Xingjia Pan, Fan Tang, Weiming Dong, Chongyang Ma, Yiping Meng, Feiyue Huang, Tong-Yee Lee, and Changsheng Xu. 2019. Content-based visual summarization for image collections. IEEE Transactions on Visualization and Computer Graphics (TVCG\u201919) (2019).","journal-title":"IEEE Transactions on Visualization and Computer Graphics (TVCG\u201919)"},{"key":"e_1_3_3_45_2","first-page":"311","volume-title":"40th Annual Meeting on Association for Computational Linguistics (ACL\u201902)","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A method for automatic evaluation of machine translation. In 40th Annual Meeting on Association for Computational Linguistics (ACL\u201902). 311\u2013318."},{"issue":"7","key":"e_1_3_3_46_2","doi-asserted-by":"crossref","first-page":"3009","DOI":"10.3390\/app11073009","article-title":"Multi view attention network for visual dialog","volume":"11","author":"Park Sungjin","year":"2021","unstructured":"Sungjin Park, Taesun Whang, Yeochan Yoon, and Heuiseok Lim. 2021. Multi view attention network for visual dialog. Applied Sciences 11, 7 (2021), 3009.","journal-title":"Applied Sciences"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_3_48_2","first-page":"779","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201916)","author":"Redmon Joseph","year":"2016","unstructured":"Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. 2016. You only look once: Unified, real time object detection. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201916). 779\u2013788."},{"key":"e_1_3_3_49_2","first-page":"3982","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP\u201919)","author":"Reimers Nils","year":"2019","unstructured":"Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP\u201919). 3982\u20133992."},{"key":"e_1_3_3_50_2","first-page":"91","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"28","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems 28 (2015), 91\u201399.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE.2007.32"},{"key":"e_1_3_3_52_2","volume-title":"ACLWeb. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing","author":"Rush Alexander M.","year":"2017","unstructured":"Alexander M. Rush, SEAS Harvard, Sumit Chopra, and Jason Weston. 2017. A neural attention model for sentence summarization. In ACLWeb. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing."},{"issue":"9","key":"e_1_3_3_53_2","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1007\/s12046-018-0908-9","article-title":"A multi-criteria context-sensitive approach for social image collection summarization","volume":"43","author":"Samani Zahra Riahi","year":"2018","unstructured":"Zahra Riahi Samani and Mohsen Ebrahimi Moghaddam. 2018. A multi-criteria context-sensitive approach for social image collection summarization. S\u0101dhan\u0101 43, 9 (2018), 143.","journal-title":"S\u0101dhan\u0101"},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475321"},{"key":"e_1_3_3_55_2","first-page":"1","volume-title":"11th IEEE\/CVF International Conference on Computer Vision (ICCV\u201907)","author":"Simon Ian","year":"2007","unstructured":"Ian Simon, Noah Snavely, and Steven M. Seitz. 2007. Scene summarization for online image collections. In 11th IEEE\/CVF International Conference on Computer Vision (ICCV\u201907). IEEE, 1\u20138."},{"key":"e_1_3_3_56_2","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. 5998\u20136008."},{"key":"e_1_3_3_57_2","first-page":"4566","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201915)","author":"Vedantam Ramakrishna","year":"2015","unstructured":"Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. CIDEr: Consensus-based image description evaluation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201915). 4566\u20134575."},{"key":"e_1_3_3_58_2","doi-asserted-by":"crossref","first-page":"169","DOI":"10.4324\/9781315095400-6","volume-title":"Copyright Law","author":"Lohmann Fred Von","year":"2017","unstructured":"Fred Von Lohmann. 2017. Fair use as innovation policy. In Copyright Law. Routledge, 169\u2013205."},{"key":"e_1_3_3_59_2","first-page":"7281","volume-title":"33rd AAAI Conference on Artificial Intelligence (AAAI\u201919)","volume":"33","author":"Wu Yu","year":"2019","unstructured":"Yu Wu, Furu Wei, Shaohan Huang, Yunli Wang, Zhoujun Li, and Ming Zhou. 2019. Response generation by context aware prototype editing. In 33rd AAAI Conference on Artificial Intelligence (AAAI\u201919), Vol. 33. 7281\u20137288."},{"key":"e_1_3_3_60_2","first-page":"3981","volume-title":"Conference on Empirical Methods in Natural Language Processing (EMNLP\u201918)","author":"Xinnuo Xu","year":"2018","unstructured":"Xu Xinnuo, Ondrej Dusek, Ioannis Konstas, and Verena Rieser. 2018. Better conversations by modeling, filtering, and optimizing for coherence and diversity. In Conference on Empirical Methods in Natural Language Processing (EMNLP\u201918). Association for Computational Linguistics, 3981\u20133991."},{"key":"e_1_3_3_61_2","article-title":"Better conversations by modeling, filtering, and optimizing for coherence and diversity","author":"Xu Xinnuo","year":"2018","unstructured":"Xinnuo Xu, Ond\u0159ej Du\u0161ek, Ioannis Konstas, and Verena Rieser. 2018. Better conversations by modeling, filtering, and optimizing for coherence and diversity. arXiv preprint arXiv:1809.06873 (2018).","journal-title":"arXiv preprint arXiv:1809.06873"},{"key":"e_1_3_3_62_2","first-page":"9360","volume-title":"Conference on Empirical Methods in Natural Language Processing (EMNLP\u201920)","author":"Zellers Rowan","year":"2020","unstructured":"Rowan Zellers, Jiasen Lu, Ximing Lu, Youngjae Yu, Yanpeng Zhao, Mohammadreza Salehi, Aditya Kusupati, Jack Hessel, Ali Farhadi, and Yejin Choi. 2020. VMSMO: Learning to generate multimodal summary for video-based news articles. In Conference on Empirical Methods in Natural Language Processing (EMNLP\u201920). 9360\u20139369."},{"key":"e_1_3_3_63_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1061"},{"key":"e_1_3_3_64_2","volume-title":"32nd AAAI Conference on Artificial Intelligence (AAAI\u201918)","author":"Zhou Hao","year":"2018","unstructured":"Hao Zhou, Minlie Huang, Tianyang Zhang, Xiaoyan Zhu, and Bing Liu. 2018. Emotional chatting machine: Emotional conversation generation with internal and external memory. In 32nd AAAI Conference on Artificial Intelligence (AAAI\u201918)."},{"key":"e_1_3_3_65_2","first-page":"1128","volume-title":"56th Annual Meeting of the Association for Computational Linguistics (ACL\u201918)","author":"Zhou Xianda","year":"2018","unstructured":"Xianda Zhou and William Yang Wang. 2018. MojiTalk: Generating emotional responses at scale. In 56th Annual Meeting of the Association for Computational Linguistics (ACL\u201918). 1128\u20131137."},{"key":"e_1_3_3_66_2","first-page":"4154","volume-title":"Conference on Empirical Methods in Natural Language Processing (EMNLP\u201918)","author":"Zhu Junnan","year":"2018","unstructured":"Junnan Zhu, Haoran Li, Tianshang Liu, Yu Zhou, Jiajun Zhang, and Chengqing Zong. 2018. MSMO: Multimodal summarization with multimodal output. In Conference on Empirical Methods in Natural Language Processing (EMNLP\u201918). 4154\u20134164."},{"key":"e_1_3_3_67_2","first-page":"9749","volume-title":"34th AAAI Conference on Artificial Intelligence (AAAI\u201920)","volume":"34","author":"Zhu Junnan","year":"2020","unstructured":"Junnan Zhu, Yu Zhou, Jiajun Zhang, Haoran Li, Chengqing Zong, and Changliang Li. 2020. Multimodal summarization with guidance of multimodal reference. In 34th AAAI Conference on Artificial Intelligence (AAAI\u201920), Vol. 34. 9749\u20139756."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589186","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3589186","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:48:53Z","timestamp":1750182533000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589186"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,31]]},"references-count":66,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1145\/3589186"],"URL":"https:\/\/doi.org\/10.1145\/3589186","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2023,5,31]]},"assertion":[{"value":"2022-03-23","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-19","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-31","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}