{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T18:39:49Z","timestamp":1784227189207,"version":"3.55.0"},"reference-count":172,"publisher":"Association for Computing Machinery (ACM)","issue":"2s","license":[{"start":{"date-parts":[[2023,2,17]],"date-time":"2023-02-17T00:00:00Z","timestamp":1676592000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Zhejiang Provincial Natural Science Foundation of China","award":["LR19F020004"],"award-info":[{"award-number":["LR19F020004"]}]},{"name":"National Key Research and Development Program of China","award":["2020AAA0107400"],"award-info":[{"award-number":["2020AAA0107400"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U20A20222"],"award-info":[{"award-number":["U20A20222"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100014219","name":"National Science Foundation for Distinguished Young Scholars","doi-asserted-by":"crossref","award":["62225605"],"award-info":[{"award-number":["62225605"]}],"id":[{"id":"10.13039\/501100014219","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100018735","name":"Ant Group","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100018735","id-type":"DOI","asserted-by":"crossref"}]},{"name":"CAAI-HUAWEI MindSpore Open Fund"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,6,30]]},"abstract":"<jats:p>Deep Learning has implemented a wide range of applications and has become increasingly popular in recent years. The goal of multimodal deep learning (MMDL) is to create models that can process and link information using various modalities. Despite the extensive development made for unimodal learning, it still cannot cover all the aspects of human learning. Multimodal learning helps to understand and analyze better when various senses are engaged in the processing of information. This article focuses on multiple types of modalities, i.e., image, video, text, audio, body gestures, facial expressions, physiological signals, flow, RGB, pose, depth, mesh, and point cloud. Detailed analysis of the baseline approaches and an in-depth study of recent advancements during the past five years (2017 to 2021) in multimodal deep learning applications has been provided. A fine-grained taxonomy of various multimodal deep learning methods is proposed, elaborating on different applications in more depth. Last, main issues are highlighted separately for each domain, along with their possible future research directions.<\/jats:p>","DOI":"10.1145\/3545572","type":"journal-article","created":{"date-parts":[[2022,10,27]],"date-time":"2022-10-27T12:29:06Z","timestamp":1666873746000},"page":"1-41","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":174,"title":["A Review on Methods and Applications in Multimodal Deep Learning"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0655-8414","authenticated-orcid":false,"given":"Summaira","family":"Jabeen","sequence":"first","affiliation":[{"name":"College of Computer Science, Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3023-1662","authenticated-orcid":false,"given":"Xi","family":"Li","sequence":"additional","affiliation":[{"name":"College of Computer Science, Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9046-1502","authenticated-orcid":false,"given":"Muhammad Shoib","family":"Amin","sequence":"additional","affiliation":[{"name":"School of Software Engineering, East China Normal University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7410-3825","authenticated-orcid":false,"given":"Omar","family":"Bourahla","sequence":"additional","affiliation":[{"name":"College of Computer Science, Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4052-1006","authenticated-orcid":false,"given":"Songyuan","family":"Li","sequence":"additional","affiliation":[{"name":"College of Computer Science, Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0309-1476","authenticated-orcid":false,"given":"Abdul","family":"Jabbar","sequence":"additional","affiliation":[{"name":"College of Computer Science, Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,2,17]]},"reference":[{"key":"e_1_3_3_2_2","first-page":"12487","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Aafaq Nayyer","year":"2019","unstructured":"Nayyer Aafaq, Naveed Akhtar, Wei Liu, Syed Zulqarnain Gilani, and Ajmal Mian. 2019. Spatio-temporal dynamics and semantic attribute enriched visual encoding for video captioning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 12487\u201312496."},{"issue":"6","key":"e_1_3_3_3_2","article-title":"VQA-Med: Overview of the medical visual question answering task at ImageCLEF 2019.","volume":"2","author":"Abacha Asma Ben","year":"2019","unstructured":"Asma Ben Abacha, Sadid A. Hasan, Vivek V. Datla, Joey Liu, Dina Demner-Fushman, and Henning M\u00fcller. 2019. VQA-Med: Overview of the medical visual question answering task at ImageCLEF 2019.CLEF (Working Notes) 2, 6 (2019).","journal-title":"CLEF (Working Notes)"},{"key":"e_1_3_3_4_2","first-page":"4971","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern recognition","author":"Agrawal Aishwarya","year":"2018","unstructured":"Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. 2018. Don\u2019t just assume; look and answer: Overcoming priors for visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern recognition. 4971\u20134980."},{"key":"e_1_3_3_5_2","first-page":"6077","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Anderson Peter","year":"2018","unstructured":"Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 6077\u20136086."},{"key":"e_1_3_3_6_2","first-page":"2425","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Antol Stanislaw","year":"2015","unstructured":"Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: Visual question answering. In Proceedings of the IEEE International Conference on Computer Vision. 2425\u20132433."},{"key":"e_1_3_3_7_2","article-title":"Neural voice cloning with a few samples","volume":"31","author":"Arik Sercan","year":"2018","unstructured":"Sercan Arik, Jitong Chen, Kainan Peng, Wei Ping, and Yanqi Zhou. 2018. Neural voice cloning with a few samples. Adv. Neural Inf. Process. Syst. 31 (2018).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_3_8_2","first-page":"195","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Ar\u0131k Sercan \u00d6.","year":"2017","unstructured":"Sercan \u00d6. Ar\u0131k, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et\u00a0al. 2017. Deep voice: Real-time neural text-to-speech. In Proceedings of the International Conference on Machine Learning. PMLR, 195\u2013204."},{"issue":"6","key":"e_1_3_3_9_2","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1007\/s00530-010-0182-0","article-title":"Multimodal fusion for multimedia analysis: A survey","volume":"16","author":"Atrey Pradeep K.","year":"2010","unstructured":"Pradeep K. Atrey, M. Anwar Hossain, Abdulmotaleb El Saddik, and Mohan S. Kankanhalli. 2010. Multimodal fusion for multimedia analysis: A survey. Multim. Syst. 16, 6 (2010), 345\u2013379.","journal-title":"Multim. Syst."},{"key":"e_1_3_3_10_2","doi-asserted-by":"crossref","first-page":"722","DOI":"10.1007\/978-3-540-76298-0_52","volume-title":"The Semantic Web","author":"Auer S\u00f6ren","year":"2007","unstructured":"S\u00f6ren Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. 2007. DBpedia: A nucleus for a web of open data. In The Semantic Web. Springer, 722\u2013735."},{"issue":"4","key":"e_1_3_3_11_2","doi-asserted-by":"crossref","first-page":"429","DOI":"10.1016\/S0163-6383(83)90241-2","article-title":"Infants\u2019 perception of substance and temporal synchrony in multimodal events","volume":"6","author":"Bahrick Lorraine E.","year":"1983","unstructured":"Lorraine E. Bahrick. 1983. Infants\u2019 perception of substance and temporal synchrony in multimodal events. Infant Behav. Devel. 6, 4 (1983), 429\u2013451.","journal-title":"Infant Behav. Devel."},{"issue":"2","key":"e_1_3_3_12_2","doi-asserted-by":"crossref","first-page":"423","DOI":"10.1109\/TPAMI.2018.2798607","article-title":"Multimodal machine learning: A survey and taxonomy","volume":"41","author":"Baltru\u0161aitis Tadas","year":"2019","unstructured":"Tadas Baltru\u0161aitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2019. Multimodal machine learning: A survey and taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 41, 2 (2019), 423\u2013443.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_3_13_2","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1007\/978-3-030-39197-3_4","volume-title":"Proceedings of the International Symposium on Practical Aspects of Declarative Languages","author":"Basu Kinjal","year":"2020","unstructured":"Kinjal Basu, Farhad Shakerin, and Gopal Gupta. 2020. AQuA: ASP-based visual question answering. In Proceedings of the International Symposium on Practical Aspects of Declarative Languages. Springer, 57\u201372."},{"key":"e_1_3_3_14_2","first-page":"1","volume-title":"Proceedings of the 1st ACM International Conference on Multimedia Retrieval","author":"Beecks Christian","year":"2011","unstructured":"Christian Beecks, Jakub Loko\u010d, Thomas Seidl, and Tom\u00e1\u0161 Skopal. 2011. Indexing the signature quadratic form distance for efficient content-based multimedia retrieval. In Proceedings of the 1st ACM International Conference on Multimedia Retrieval. 1\u20138."},{"key":"e_1_3_3_15_2","first-page":"2612","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Ben-Younes Hedi","year":"2017","unstructured":"Hedi Ben-Younes, R\u00e9mi Cadene, Matthieu Cord, and Nicolas Thome. 2017. MUTAN: Multimodal Tucker fusion for visual question answering. In Proceedings of the IEEE International Conference on Computer Vision. 2612\u20132620."},{"key":"e_1_3_3_16_2","article-title":"Experience grounds language","author":"Bisk Yonatan","year":"2020","unstructured":"Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, et\u00a0al. 2020. Experience grounds language. arXiv preprint arXiv:2004.10151 (2020).","journal-title":"arXiv preprint arXiv:2004.10151"},{"key":"e_1_3_3_17_2","first-page":"748","volume-title":"Proceedings of the International Conference on Multimodal Interaction","author":"Boateng George","year":"2020","unstructured":"George Boateng. 2020. Towards real-time multimodal emotion recognition among couples. In Proceedings of the International Conference on Multimodal Interaction. 748\u2013753."},{"key":"e_1_3_3_18_2","first-page":"1247","volume-title":"Proceedings of the ACM SIGMOD International Conference on Management of Data","author":"Bollacker Kurt","year":"2008","unstructured":"Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor. 2008. Freebase: A collaboratively created graph database for structuring human knowledge. In Proceedings of the ACM SIGMOD International Conference on Management of Data. 1247\u20131250."},{"issue":"4","key":"e_1_3_3_19_2","doi-asserted-by":"crossref","first-page":"335","DOI":"10.1007\/s10579-008-9076-6","article-title":"IEMOCAP: Interactive emotional dyadic motion capture database","volume":"42","author":"Busso Carlos","year":"2008","unstructured":"Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan. 2008. IEMOCAP: Interactive emotional dyadic motion capture database. Lang. Resour. Eval. 42, 4 (2008), 335\u2013359.","journal-title":"Lang. Resour. Eval."},{"issue":"1","key":"e_1_3_3_20_2","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1109\/TAFFC.2016.2515617","article-title":"MSP-IMPROV: An acted corpus of dyadic interactions to study emotion perception","volume":"8","author":"Busso Carlos","year":"2016","unstructured":"Carlos Busso, Srinivas Parthasarathy, Alec Burmania, Mohammed AbdelWahab, Najmeh Sadoughi, and Emily Mower Provost. 2016. MSP-IMPROV: An acted corpus of dyadic interactions to study emotion perception. IEEE Trans. Affect. Comput. 8, 1 (2016), 67\u201380.","journal-title":"IEEE Trans. Affect. Comput."},{"key":"e_1_3_3_21_2","first-page":"961","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Heilbron Fabian Caba","year":"2015","unstructured":"Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015. ActivityNet: A large-scale video benchmark for human activity understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 961\u2013970."},{"key":"e_1_3_3_22_2","first-page":"1989","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Cadene Remi","year":"2019","unstructured":"Remi Cadene, Hedi Ben-Younes, Matthieu Cord, and Nicolas Thome. 2019. MUREL: Multimodal relational reasoning for visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 1989\u20131998."},{"issue":"1","key":"e_1_3_3_23_2","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1007\/s11063-018-09973-5","article-title":"Image captioning with bidirectional semantic attention-based guiding of long short-term memory","volume":"50","author":"Cao Pengfei","year":"2019","unstructured":"Pengfei Cao, Zhongyi Yang, Liang Sun, Yanchun Liang, Mary Qu Yang, and Renchu Guan. 2019. Image captioning with bidirectional semantic attention-based guiding of long short-term memory. Neural Process. Lett. 50, 1 (2019), 103\u2013119.","journal-title":"Neural Process. Lett."},{"key":"e_1_3_3_24_2","first-page":"28","volume-title":"Proceedings of the International Workshop on Machine Learning for Multimodal Interaction","author":"Carletta Jean","year":"2005","unstructured":"Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, et\u00a0al. 2005. The AMI meeting corpus: A pre-announcement. In Proceedings of the International Workshop on Machine Learning for Multimodal Interaction. Springer, 28\u201339."},{"key":"e_1_3_3_25_2","first-page":"190","volume-title":"Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies","author":"Chen David","year":"2011","unstructured":"David Chen and William B. Dolan. 2011. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies. 190\u2013200."},{"issue":"14","key":"e_1_3_3_26_2","doi-asserted-by":"crossref","first-page":"8669","DOI":"10.1007\/s00521-020-05616-w","article-title":"HEU Emotion: A large-scale database for multimodal emotion recognition in the wild","volume":"33","author":"Chen Jing","year":"2021","unstructured":"Jing Chen, Chenhui Wang, Kejun Wang, Chaoqun Yin, Cong Zhao, Tao Xu, Xinyi Zhang, Ziqiang Huang, Meichen Liu, and Tao Yang. 2021. HEU Emotion: A large-scale database for multimodal emotion recognition in the wild. Neural Comput. Applic. 33, 14 (2021), 8669\u20138685.","journal-title":"Neural Comput. Applic."},{"key":"e_1_3_3_27_2","first-page":"16846","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen Long","year":"2021","unstructured":"Long Chen, Zhihong Jiang, Jun Xiao, and Wei Liu. 2021. Human-like controllable image captioning with verb-specific semantic roles. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 16846\u201316856."},{"key":"e_1_3_3_28_2","volume-title":"Proceedings of the 31st AAAI Conference on Artificial Intelligence","author":"Chen Minghai","year":"2017","unstructured":"Minghai Chen, Guiguang Ding, Sicheng Zhao, Hui Chen, Qiang Liu, and Jungong Han. 2017. Reference-based LSTM for image captioning. In Proceedings of the 31st AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_3_29_2","article-title":"Microsoft COCO captions: Data collection and evaluation server","author":"Chen Xinlei","year":"2015","unstructured":"Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2015. Microsoft COCO captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325 (2015).","journal-title":"arXiv preprint arXiv:1504.00325"},{"key":"e_1_3_3_30_2","first-page":"358","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Chen Yangyu","year":"2018","unstructured":"Yangyu Chen, Shuhui Wang, Weigang Zhang, and Qingming Huang. 2018. Less is more: Picking informative frames for video captioning. In Proceedings of the European Conference on Computer Vision (ECCV). 358\u2013373."},{"key":"e_1_3_3_31_2","doi-asserted-by":"crossref","first-page":"154953","DOI":"10.1109\/ACCESS.2020.3018752","article-title":"Stack-VS: Stacked visual-semantic attention for image caption generation","volume":"8","author":"Cheng Ling","year":"2020","unstructured":"Ling Cheng, Wei Wei, Xianling Mao, Yong Liu, and Chunyan Miao. 2020. Stack-VS: Stacked visual-semantic attention for image caption generation. IEEE Access 8 (2020), 154953\u2013154965.","journal-title":"IEEE Access"},{"key":"e_1_3_3_32_2","first-page":"213","volume-title":"Proceedings of the 5th International Conference on Big Data Computing and Communications (BIGCOM)","author":"Chong Luyao","year":"2019","unstructured":"Luyao Chong, Meng Jin, and Yuan He. 2019. EmoChat: Bringing multimodal emotion detection to mobile conversation. In Proceedings of the 5th International Conference on Big Data Computing and Communications (BIGCOM). IEEE, 213\u2013221."},{"key":"e_1_3_3_33_2","doi-asserted-by":"crossref","first-page":"168865","DOI":"10.1109\/ACCESS.2020.3023871","article-title":"Cross-subject multimodal emotion recognition based on hybrid fusion","volume":"8","author":"Cimtay Yucel","year":"2020","unstructured":"Yucel Cimtay, Erhan Ekmekcioglu, and Seyma Caglar-Ozhan. 2020. Cross-subject multimodal emotion recognition based on hybrid fusion. IEEE Access 8 (2020), 168865\u2013168878.","journal-title":"IEEE Access"},{"key":"e_1_3_3_34_2","doi-asserted-by":"crossref","unstructured":"Mutlu Cukurova Michail Giannakos and Roberto Martinez-Maldonado. 2020. The promise and challenges of multimodal learning analytics. (2020) 1441\u20131449 pages.","DOI":"10.1111\/bjet.13015"},{"issue":"4","key":"e_1_3_3_35_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3103613","article-title":"Multimodal retrieval with diversification and relevance feedback for tourist attraction images","volume":"13","author":"Dang-Nguyen Duc-Tien","year":"2017","unstructured":"Duc-Tien Dang-Nguyen, Luca Piras, Giorgio Giacinto, Giulia Boato, and Francesco GB DE Natale. 2017. Multimodal retrieval with diversification and relevance feedback for tourist attraction images. ACM Trans. Multim. Comput. Commun. Applic. 13, 4 (2017), 1\u201324.","journal-title":"ACM Trans. Multim. Comput. Commun. Applic."},{"key":"e_1_3_3_36_2","first-page":"1814","volume-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Desta Mikyas T.","year":"2018","unstructured":"Mikyas T. Desta, Larry Chen, and Tomasz Kornuta. 2018. Object-based reasoning in VQA. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 1814\u20131823."},{"key":"e_1_3_3_37_2","volume-title":"Proceedings of the International Conference on Innovative Computing & Communication (ICICC)","author":"Diwakar Parul","year":"2021","unstructured":"Parul Diwakar. 2021. Automatic image captioning using deep learning. In Proceedings of the International Conference on Innovative Computing & Communication (ICICC)."},{"key":"e_1_3_3_38_2","first-page":"5709","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Elias Isaac","year":"2021","unstructured":"Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang, Ye Jia, Ron J. Weiss, and Yonghui Wu. 2021. Parallel Tacotron: Non-autoregressive and controllable TTS. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 5709\u20135713."},{"key":"e_1_3_3_39_2","article-title":"Video2Commonsense: Generating commonsense descriptions to enrich video captioning","author":"Fang Zhiyuan","year":"2020","unstructured":"Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral, and Yezhou Yang. 2020. Video2Commonsense: Generating commonsense descriptions to enrich video captioning. arXiv preprint arXiv:2003.05162 (2020).","journal-title":"arXiv preprint arXiv:2003.05162"},{"key":"e_1_3_3_40_2","first-page":"4125","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Feng Yang","year":"2019","unstructured":"Yang Feng, Lin Ma, Wei Liu, and Jiebo Luo. 2019. Unsupervised image captioning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 4125\u20134134."},{"key":"e_1_3_3_41_2","first-page":"3137","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Gan Chuang","year":"2017","unstructured":"Chuang Gan, Zhe Gan, Xiaodong He, Jianfeng Gao, and Li Deng. 2017. StyleNet: Generating attractive visual captions with styles. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3137\u20133146."},{"issue":"5","key":"e_1_3_3_42_2","doi-asserted-by":"crossref","first-page":"829","DOI":"10.1162\/neco_a_01273","article-title":"A survey on deep learning for multimodal data fusion","volume":"32","author":"Gao Jing","year":"2020","unstructured":"Jing Gao, Peng Li, Zhikui Chen, and Jianing Zhang. 2020. A survey on deep learning for multimodal data fusion. Neural Computat. 32, 5 (2020), 829\u2013864.","journal-title":"Neural Computat."},{"key":"e_1_3_3_43_2","first-page":"324","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Gao Ruohan","year":"2019","unstructured":"Ruohan Gao and Kristen Grauman. 2019. 2.5D visual sound. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 324\u2013333."},{"issue":"3","key":"e_1_3_3_44_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2967502","article-title":"Event classification in microblogs via social tracking","volume":"8","author":"Gao Yue","year":"2017","unstructured":"Yue Gao, Hanwang Zhang, Xibin Zhao, and Shuicheng Yan. 2017. Event classification in microblogs via social tracking. ACM Trans. Intell. Syst. Technol. 8, 3 (2017), 1\u201314.","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"e_1_3_3_45_2","first-page":"2755","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Garcia Nuno Cruz","year":"2021","unstructured":"Nuno Cruz Garcia, Sarah Adel Bargal, Vitaly Ablavsky, Pietro Morerio, Vittorio Murino, and Stan Sclaroff. 2021. Distillation multiple choice learning for multimodal action recognition. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 2755\u20132764."},{"key":"e_1_3_3_46_2","article-title":"Deep Voice 2: Multi-speaker neural text-to-speech","volume":"30","author":"Gibiansky Andrew","year":"2017","unstructured":"Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou. 2017. Deep Voice 2: Multi-speaker neural text-to-speech. Adv. Neural Inf. Process. Syst. 30 (2017).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_3_47_2","first-page":"1148","volume-title":"Proceedings of the 18th International Conference on Pattern Recognition (ICPR\u201906)","author":"Gunes Hatice","year":"2006","unstructured":"Hatice Gunes and Massimo Piccardi. 2006. A bimodal face and body gesture database for automatic analysis of human nonverbal affective behavior. In Proceedings of the 18th International Conference on Pattern Recognition (ICPR\u201906). IEEE, 1148\u20131153."},{"key":"e_1_3_3_48_2","first-page":"4204","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Guo Longteng","year":"2019","unstructured":"Longteng Guo, Jing Liu, Peng Yao, Jiangwei Li, and Hanqing Lu. 2019. MSCap: Multi-style image captioning with unpaired stylized text. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 4204\u20134213."},{"key":"e_1_3_3_49_2","doi-asserted-by":"crossref","first-page":"63373","DOI":"10.1109\/ACCESS.2019.2916887","article-title":"Deep multimodal representation learning: A survey","volume":"7","author":"Guo Wenzhong","year":"2019","unstructured":"Wenzhong Guo, Jianwen Wang, and Shiping Wang. 2019. Deep multimodal representation learning: A survey. IEEE Access 7 (2019), 63373\u201363394.","journal-title":"IEEE Access"},{"key":"e_1_3_3_50_2","doi-asserted-by":"crossref","first-page":"6730","DOI":"10.1109\/TIP.2021.3097180","article-title":"Re-attention for visual question answering","volume":"30","author":"Guo Wenya","year":"2021","unstructured":"Wenya Guo, Ying Zhang, Jufeng Yang, and Xiaojie Yuan. 2021. Re-attention for visual question answering. IEEE Trans. Image Process. 30 (2021), 6730\u20136743.","journal-title":"IEEE Trans. Image Process."},{"key":"e_1_3_3_51_2","first-page":"3608","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Gurari Danna","year":"2018","unstructured":"Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. 2018. VizWiz grand challenge: Answering visual questions from blind people. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3608\u20133617."},{"key":"e_1_3_3_52_2","first-page":"1","volume-title":"Proceedings of the SIGGRAPH Asia Technical Communications Conference","author":"Hao Jiaqi","year":"2021","unstructured":"Jiaqi Hao, Shiguang Liu, and Qing Xu. 2021. Controlling eye blink for talking face generation via eye conversion. In Proceedings of the SIGGRAPH Asia Technical Communications Conference. 1\u20134."},{"key":"e_1_3_3_53_2","first-page":"196","volume-title":"Proceedings of the IEEE Conference on Multimedia Information Processing and Retrieval (MIPR)","author":"Hazarika Devamanyu","year":"2018","unstructured":"Devamanyu Hazarika, Sruthi Gorantla, Soujanya Poria, and Roger Zimmermann. 2018. Self-attentive feature-level fusion for multimodal emotion detection. In Proceedings of the IEEE Conference on Multimedia Information Processing and Retrieval (MIPR). IEEE, 196\u2013201."},{"key":"e_1_3_3_54_2","first-page":"2594","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Hazarika Devamanyu","year":"2018","unstructured":"Devamanyu Hazarika, Soujanya Poria, Rada Mihalcea, Erik Cambria, and Roger Zimmermann. 2018. ICON: Interactive conversational memory network for multimodal emotion detection. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. 2594\u20132604."},{"key":"e_1_3_3_55_2","doi-asserted-by":"crossref","first-page":"229","DOI":"10.1016\/j.patrec.2017.10.018","article-title":"Image caption generation with part of speech guidance","volume":"119","author":"He Xinwei","year":"2019","unstructured":"Xinwei He, Baoguang Shi, Xiang Bai, Gui-Song Xia, Zhaoxiang Zhang, and Weisheng Dong. 2019. Image caption generation with part of speech guidance. Pattern Recog. Lett. 119 (2019), 229\u2013237.","journal-title":"Pattern Recog. Lett."},{"key":"e_1_3_3_56_2","doi-asserted-by":"crossref","first-page":"853","DOI":"10.1613\/jair.3994","article-title":"Framing image description as a ranking task: Data, models and evaluation metrics","volume":"47","author":"Hodosh Micah","year":"2013","unstructured":"Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013. Framing image description as a ranking task: Data, models and evaluation metrics. J. Artif. Intell. Res. 47 (2013), 853\u2013899.","journal-title":"J. Artif. Intell. Res."},{"key":"e_1_3_3_57_2","doi-asserted-by":"crossref","first-page":"794","DOI":"10.2307\/1130130","article-title":"A multimodal assessment of behavioral and cognitive deficits in abused and neglected preschoolers","author":"Hoffman-Plotkin Debbie","year":"1984","unstructured":"Debbie Hoffman-Plotkin and Craig T. Twentyman. 1984. A multimodal assessment of behavioral and cognitive deficits in abused and neglected preschoolers. Child Devel. 55, 3 (1984), 794\u2013802.","journal-title":"Child Devel."},{"issue":"5","key":"e_1_3_3_58_2","doi-asserted-by":"crossref","first-page":"4340","DOI":"10.1109\/TGRS.2020.3016820","article-title":"More diverse means better: Multimodal deep learning meets remote-sensing imagery classification","volume":"59","author":"Hong Danfeng","year":"2020","unstructured":"Danfeng Hong, Lianru Gao, Naoto Yokoya, Jing Yao, Jocelyn Chanussot, Qian Du, and Bing Zhang. 2020. More diverse means better: Multimodal deep learning meets remote-sensing imagery classification. IEEE Trans. Geosci. Rem. Sens. 59, 5 (2020), 4340\u20134354.","journal-title":"IEEE Trans. Geosci. Rem. Sens."},{"issue":"6","key":"e_1_3_3_59_2","doi-asserted-by":"crossref","first-page":"8213","DOI":"10.1007\/s11042-020-10030-4","article-title":"Video multimodal emotion recognition based on Bi-GRU and attention fusion","volume":"80","author":"Huan Ruo-Hong","year":"2021","unstructured":"Ruo-Hong Huan, Jia Shu, Sheng-Lin Bao, Rong-Hua Liang, Peng Chen, and Kai-Kai Chi. 2021. Video multimodal emotion recognition based on Bi-GRU and attention fusion. Multim. Tools. Applic. 80, 6 (2021), 8213\u20138240.","journal-title":"Multim. Tools. Applic."},{"key":"e_1_3_3_60_2","article-title":"Learning multimodal deep representations for crowd anomaly event detection","volume":"2018","author":"Huang Shaonian","year":"2018","unstructured":"Shaonian Huang, Dongjun Huang, and Xinmin Zhou. 2018. Learning multimodal deep representations for crowd anomaly event detection. Math. Prob. Eng. 2018 (2018).","journal-title":"Math. Prob. Eng."},{"key":"e_1_3_3_61_2","article-title":"Fusion of facial expressions and EEG for multimodal emotion recognition","volume":"2017","author":"Huang Yongrui","year":"2017","unstructured":"Yongrui Huang, Jianhao Yang, Pengkai Liao, and Jiahui Pan. 2017. Fusion of facial expressions and EEG for multimodal emotion recognition. Computat. Intell. Neurosci. 2017 (2017).","journal-title":"Computat. Intell. Neurosci."},{"issue":"4","key":"e_1_3_3_62_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3409332","article-title":"Knowledge-driven egocentric multimodal activity recognition","volume":"16","author":"Huang Yi","year":"2020","unstructured":"Yi Huang, Xiaoshan Yang, Junyu Gao, Jitao Sang, and Changsheng Xu. 2020. Knowledge-driven egocentric multimodal activity recognition. ACM Trans. Multim. Comput. Commun. Applic. 16, 4 (2020), 1\u2013133.","journal-title":"ACM Trans. Multim. Comput. Commun. Applic."},{"key":"e_1_3_3_63_2","unstructured":"Keith Ito and Linda Johnson. 2017. The LJ speech dataset. Retrieved from https:\/\/keithito.com\/LJ-Speech-Dataset."},{"key":"e_1_3_3_64_2","first-page":"7415","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Jaiswal Mimansa","year":"2019","unstructured":"Mimansa Jaiswal, Zakaria Aldeneh, Cristian-Paul Bara, Yuanhang Luo, Mihai Burzo, Rada Mihalcea, and Emily Mower Provost. 2019. Muse-ing on the impact of utterance ordering on crowdsourced emotion annotations. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 7415\u20137419."},{"key":"e_1_3_3_65_2","first-page":"174","volume-title":"Proceedings of the International Conference on Multimodal Interaction","author":"Jaiswal Mimansa","year":"2019","unstructured":"Mimansa Jaiswal, Zakaria Aldeneh, and Emily Mower Provost. 2019. Controlling for confounders in multimodal emotion classification via adversarial learning. In Proceedings of the International Conference on Multimodal Interaction. 174\u2013184."},{"key":"e_1_3_3_66_2","first-page":"1655","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Ji Jiayi","year":"2021","unstructured":"Jiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen, Gen Luo, Yongjian Wu, Yue Gao, and Rongrong Ji. 2021. Improving image captioning by leveraging intra-and inter-layer global representation in transformer network. In Proceedings of the AAAI Conference on Artificial Intelligence. 1655\u20131663."},{"key":"e_1_3_3_67_2","doi-asserted-by":"crossref","first-page":"69700","DOI":"10.1109\/ACCESS.2021.3067607","article-title":"Multi-gate attention network for image captioning","volume":"9","author":"Jiang Weitao","year":"2021","unstructured":"Weitao Jiang, Xiying Li, Haifeng Hu, Qiang Lu, and Bohong Liu. 2021. Multi-gate attention network for image captioning. IEEE Access 9 (2021), 69700\u201369709.","journal-title":"IEEE Access"},{"key":"e_1_3_3_68_2","first-page":"499","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Jiang Wenhao","year":"2018","unstructured":"Wenhao Jiang, Lin Ma, Yu-Gang Jiang, Wei Liu, and Tong Zhang. 2018. Recurrent fusion network for image captioning. In Proceedings of the European Conference on Computer Vision (ECCV). 499\u2013515."},{"key":"e_1_3_3_69_2","first-page":"3142","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Jing Longlong","year":"2021","unstructured":"Longlong Jing, Elahe Vahdani, Jiaxing Tan, and Yingli Tian. 2021. Cross-modal center loss for 3D cross-modal retrieval. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3142\u20133151."},{"key":"e_1_3_3_70_2","first-page":"2901","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Johnson Justin","year":"2017","unstructured":"Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017. CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2901\u20132910."},{"issue":"3","key":"e_1_3_3_71_2","doi-asserted-by":"crossref","first-page":"251","DOI":"10.1080\/00401706.1991.10484833","article-title":"Hidden Markov models for speech recognition","volume":"33","author":"Juang Biing Hwang","year":"1991","unstructured":"Biing Hwang Juang and Laurence R. Rabiner. 1991. Hidden Markov models for speech recognition. Technometrics 33, 3 (1991), 251\u2013272.","journal-title":"Technometrics"},{"key":"e_1_3_3_72_2","first-page":"1965","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Kafle Kushal","year":"2017","unstructured":"Kushal Kafle and Christopher Kanan. 2017. An analysis of visual question answering algorithms. In Proceedings of the IEEE International Conference on Computer Vision. 1965\u20131973."},{"key":"e_1_3_3_73_2","first-page":"5492","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Kazakos Evangelos","year":"2019","unstructured":"Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen. 2019. Epic-fusion: Audio-visual temporal binding for egocentric action recognition. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 5492\u20135501."},{"key":"e_1_3_3_74_2","first-page":"1","volume-title":"Proceedings of the IEEE 13th Image, Video, and Multidimensional Signal Processing Workshop (IVMSP)","author":"Koutras Petros","year":"2018","unstructured":"Petros Koutras, Athanasia Zlatinsi, and Petros Maragos. 2018. Exploring CNN-based architectures for multimodal salient event detection in videos. In Proceedings of the IEEE 13th Image, Video, and Multidimensional Signal Processing Workshop (IVMSP). IEEE, 1\u20135."},{"key":"e_1_3_3_75_2","first-page":"706","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Krishna Ranjay","year":"2017","unstructured":"Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles. 2017. Dense-captioning events in videos. In Proceedings of the IEEE International Conference on Computer Vision. 706\u2013715."},{"issue":"1","key":"e_1_3_3_76_2","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1007\/s11263-016-0981-7","article-title":"Visual genome: Connecting language and vision using crowdsourced dense image annotations","volume":"123","author":"Krishna Ranjay","year":"2017","unstructured":"Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, et\u00a0al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vis. 123, 1 (2017), 32\u201373.","journal-title":"Int. J. Comput. Vis."},{"key":"e_1_3_3_77_2","doi-asserted-by":"crossref","first-page":"119516","DOI":"10.1109\/ACCESS.2020.3005664","article-title":"Different contextual window sizes based RNNs for multimodal emotion detection in interactive conversations","volume":"8","author":"Lai Helang","year":"2020","unstructured":"Helang Lai, Hongying Chen, and Shuangyan Wu. 2020. Different contextual window sizes based RNNs for multimodal emotion detection in interactive conversations. IEEE Access 8 (2020), 119516\u2013119526.","journal-title":"IEEE Access"},{"key":"e_1_3_3_78_2","doi-asserted-by":"crossref","DOI":"10.1097\/00005053-197306000-00005","article-title":"Multimodal behavior therapy: Treating the \u201cBASIC ID\u201d","author":"Lazarus Arnold A.","year":"1973","unstructured":"Arnold A. Lazarus. 1973. Multimodal behavior therapy: Treating the \u201cBASIC ID\u201d. J. Nerv. Ment. Dis. 156, 6 (1973).","journal-title":"J. Nerv. Ment. Dis."},{"issue":"2","key":"e_1_3_3_79_2","doi-asserted-by":"crossref","first-page":"55","DOI":"10.3390\/fi13020055","article-title":"Video captioning based on channel soft attention and semantic reconstructor","volume":"13","author":"Lei Zhou","year":"2021","unstructured":"Zhou Lei and Yiyong Huang. 2021. Video captioning based on channel soft attention and semantic reconstructor. Fut. Internet 13, 2 (2021), 55.","journal-title":"Fut. Internet"},{"key":"e_1_3_3_80_2","first-page":"10313","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Li Linjie","year":"2019","unstructured":"Linjie Li, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019. Relation-aware graph attention network for visual question answering. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 10313\u201310322."},{"key":"e_1_3_3_81_2","first-page":"339","volume-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Li Lijun","year":"2019","unstructured":"Lijun Li and Boqing Gong. 2019. End-to-end video captioning with multitask reinforcement learning. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 339\u2013348."},{"issue":"3","key":"e_1_3_3_82_2","first-page":"726","article-title":"GLA: Global\u2013local attention for image description","volume":"20","author":"Li Linghui","year":"2017","unstructured":"Linghui Li, Sheng Tang, Yongdong Zhang, Lixi Deng, and Qi Tian. 2017. GLA: Global\u2013local attention for image description. IEEE Trans. Multim. 20, 3 (2017), 726\u2013737.","journal-title":"IEEE Trans. Multim."},{"key":"e_1_3_3_83_2","doi-asserted-by":"crossref","first-page":"187208","DOI":"10.1109\/ACCESS.2020.3029288","article-title":"Multistep deep system for multimodal emotion detection with invalid data in the Internet of Things","volume":"8","author":"Li Minjia","year":"2020","unstructured":"Minjia Li, Lun Xie, Zeping Lv, Juan Li, and Zhiliang Wang. 2020. Multistep deep system for multimodal emotion detection with invalid data in the Internet of Things. IEEE Access 8 (2020), 187208\u2013187221.","journal-title":"IEEE Access"},{"key":"e_1_3_3_84_2","first-page":"271","volume-title":"Proceedings of the ACM International Conference on Multimedia Retrieval","author":"Li Xirong","year":"2016","unstructured":"Xirong Li, Weiyu Lan, Jianfeng Dong, and Hailong Liu. 2016. Adding Chinese captions to images. In Proceedings of the ACM International Conference on Multimedia Retrieval. 271\u2013275."},{"issue":"10","key":"e_1_3_3_85_2","doi-asserted-by":"crossref","first-page":"1863","DOI":"10.1109\/TKDE.2018.2872063","article-title":"A survey of multi-view representation learning","volume":"31","author":"Li Yingming","year":"2019","unstructured":"Yingming Li, Ming Yang, and Zhongfei Zhang. 2019. A survey of multi-view representation learning. IEEE Trans. Knowl. Data Eng. 31, 10 (2019), 1863\u20131883.","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"e_1_3_3_86_2","first-page":"740","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In Proceedings of the European Conference on Computer Vision. Springer, 740\u2013755."},{"issue":"4","key":"e_1_3_3_87_2","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1023\/B:BTTJ.0000047600.45421.6d","article-title":"ConceptNet\u2014a practical commonsense reasoning tool-kit","volume":"22","author":"Liu Hugo","year":"2004","unstructured":"Hugo Liu and Push Singh. 2004. ConceptNet\u2014a practical commonsense reasoning tool-kit. BT Technol. J. 22, 4 (2004), 211\u2013226.","journal-title":"BT Technol. J."},{"key":"e_1_3_3_88_2","article-title":"Chinese image caption generation via visual attention and topic modeling","author":"Liu Maofu","year":"2020","unstructured":"Maofu Liu, Huijun Hu, Lingjun Li, Yan Yu, and Weili Guan. 2020. Chinese image caption generation via visual attention and topic modeling. IEEE Trans. Cyber. 52, 2 (2020).","journal-title":"IEEE Trans. Cyber."},{"issue":"2","key":"e_1_3_3_89_2","doi-asserted-by":"crossref","first-page":"102178","DOI":"10.1016\/j.ipm.2019.102178","article-title":"Image caption generation with dual attention mechanism","volume":"57","author":"Liu Maofu","year":"2020","unstructured":"Maofu Liu, Lingjun Li, Huijun Hu, Weili Guan, and Jing Tian. 2020. Image caption generation with dual attention mechanism. Inf. Process. Manag. 57, 2 (2020), 102178.","journal-title":"Inf. Process. Manag."},{"key":"e_1_3_3_90_2","first-page":"1425","volume-title":"Proceedings of the 26th ACM International Conference on Multimedia","author":"Liu Sheng","year":"2018","unstructured":"Sheng Liu, Zhou Ren, and Junsong Yuan. 2018. SibNet: Sibling convolutional encoder for video captioning. In Proceedings of the 26th ACM International Conference on Multimedia. 1425\u20131434."},{"issue":"12","key":"e_1_3_3_91_2","doi-asserted-by":"crossref","first-page":"8555","DOI":"10.1109\/TGRS.2020.2988782","article-title":"RSVQA: Visual question answering for remote sensing data","volume":"58","author":"Lobry Sylvain","year":"2020","unstructured":"Sylvain Lobry, Diego Marcos, Jesse Murray, and Devis Tuia. 2020. RSVQA: Visual question answering for remote sensing data. IEEE Trans. Geosci. Rem. Sens. 58, 12 (2020), 8555\u20138566.","journal-title":"IEEE Trans. Geosci. Rem. Sens."},{"issue":"20","key":"e_1_3_3_92_2","doi-asserted-by":"crossref","first-page":"758","DOI":"10.1049\/ell2.12255","article-title":"Improving reasoning with contrastive visual information for visual question answering","volume":"57","author":"Long Yu","year":"2021","unstructured":"Yu Long, Pengjie Tang, Hanli Wang, and Jian Yu. 2021. Improving reasoning with contrastive visual information for visual question answering. Electron. Lett. 57, 20 (2021), 758\u2013760.","journal-title":"Electron. Lett."},{"key":"e_1_3_3_93_2","article-title":"A multi-world approach to question answering about real-world scenes based on uncertain input","volume":"27","author":"Malinowski Mateusz","year":"2014","unstructured":"Mateusz Malinowski and Mario Fritz. 2014. A multi-world approach to question answering about real-world scenes based on uncertain input. Adv. Neural Inf. Process. Syst. 27 (2014).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_3_94_2","first-page":"3195","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Marino Kenneth","year":"2019","unstructured":"Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019. OK-VQA: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3195\u20133204."},{"key":"e_1_3_3_95_2","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1109\/ICDEW.2006.145","volume-title":"Proceedings of the 22nd International Conference on Data Engineering Workshops (ICDEW\u201906)","author":"Martin Olivier","year":"2006","unstructured":"Olivier Martin, Irene Kotsia, Benoit Macq, and Ioannis Pitas. 2006. The eNTERFACE\u201905 audio-visual emotion database. In Proceedings of the 22nd International Conference on Data Engineering Workshops (ICDEW\u201906). IEEE, 8\u20138."},{"key":"e_1_3_3_96_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Mathews Alexander","year":"2016","unstructured":"Alexander Mathews, Lexing Xie, and Xuming He. 2016. SentiCap: Generating image descriptions with sentiments. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"issue":"5588","key":"e_1_3_3_97_2","doi-asserted-by":"crossref","first-page":"746","DOI":"10.1038\/264746a0","article-title":"Hearing lips and seeing voices","volume":"264","author":"McGurk Harry","year":"1976","unstructured":"Harry McGurk and John MacDonald. 1976. Hearing lips and seeing voices. Nature 264, 5588 (1976), 746\u2013748.","journal-title":"Nature"},{"issue":"1","key":"e_1_3_3_98_2","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1109\/T-AFFC.2011.20","article-title":"The SEMAINE database: Annotated multimodal records of emotionally colored conversations between a person and a limited agent","volume":"3","author":"McKeown Gary","year":"2011","unstructured":"Gary McKeown, Michel Valstar, Roddy Cowie, Maja Pantic, and Marc Schroder. 2011. The SEMAINE database: Annotated multimodal records of emotionally colored conversations between a person and a limited agent. IEEE Trans. Affect. Comput. 3, 1 (2011), 5\u201317.","journal-title":"IEEE Trans. Affect. Comput."},{"issue":"11","key":"e_1_3_3_99_2","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1145\/219717.219748","article-title":"WordNet: A lexical database for English","volume":"38","author":"Miller George A.","year":"1995","unstructured":"George A. Miller. 1995. WordNet: A lexical database for English. Commun. ACM. 38, 11 (1995), 39\u201341.","journal-title":"Commun. ACM."},{"key":"e_1_3_3_100_2","first-page":"1359","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Mittal Trisha","year":"2020","unstructured":"Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, and Dinesh Manocha. 2020. M3er: Multiplicative multimodal emotion recognition using facial, textual, and speech cues. In Proceedings of the AAAI Conference on Artificial Intelligence. 1359\u20131367."},{"key":"e_1_3_3_101_2","doi-asserted-by":"crossref","first-page":"1183","DOI":"10.1613\/jair.1.11688","article-title":"Trends in integration of vision and language research: A survey of tasks, datasets, and methods","volume":"71","author":"Mogadala Aditya","year":"2021","unstructured":"Aditya Mogadala, Marimuthu Kalimuthu, and Dietrich Klakow. 2021. Trends in integration of vision and language research: A survey of tasks, datasets, and methods. J. Artif. Intell. Res. 71 (2021), 1183\u20131317.","journal-title":"J. Artif. Intell. Res."},{"key":"e_1_3_3_102_2","unstructured":"Louis-Philippe Morency. 2020. Multimodal Machine Learning (or Deep Learning for Multimodal Systems). Retrieved from https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2017\/07\/Integrative_AI_Louis_Philippe_Morency.pdf."},{"issue":"5","key":"e_1_3_3_103_2","doi-asserted-by":"crossref","first-page":"471","DOI":"10.3758\/BF03204892","article-title":"Multimodal signal detection: Independent decisions vs. integration","volume":"28","author":"Mulligan Robert M.","year":"1980","unstructured":"Robert M. Mulligan and Marilyn L. Shaw. 1980. Multimodal signal detection: Independent decisions vs. integration. Percept. Psychophys. 28, 5 (1980), 471\u2013478.","journal-title":"Percept. Psychophys."},{"key":"e_1_3_3_104_2","first-page":"6588","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Mun Jonghwan","year":"2019","unstructured":"Jonghwan Mun, Linjie Yang, Zhou Ren, Ning Xu, and Bohyung Han. 2019. Streamlined dense video captioning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6588\u20136597."},{"key":"e_1_3_3_105_2","first-page":"451","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Narasimhan Medhini","year":"2018","unstructured":"Medhini Narasimhan and Alexander G. Schwing. 2018. Straight to the facts: Learning knowledge base retrieval for factual visual question answering. In Proceedings of the European Conference on Computer Vision (ECCV). 451\u2013468."},{"key":"e_1_3_3_106_2","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1016\/j.cviu.2018.06.005","article-title":"Deep spatio-temporal feature fusion with compact bilinear pooling for multimodal emotion recognition","volume":"174","author":"Nguyen Dung","year":"2018","unstructured":"Dung Nguyen, Kien Nguyen, Sridha Sridharan, David Dean, and Clinton Fookes. 2018. Deep spatio-temporal feature fusion with compact bilinear pooling for multimodal emotion recognition. Comput. Vis. Image Underst. 174 (2018), 33\u201342.","journal-title":"Comput. Vis. Image Underst."},{"key":"e_1_3_3_107_2","first-page":"1215","volume-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Nguyen Dung","year":"2017","unstructured":"Dung Nguyen, Kien Nguyen, Sridha Sridharan, Afsane Ghasemi, David Dean, and Clinton Fookes. 2017. Deep spatio-temporal features for multimodal emotion recognition. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 1215\u20131223."},{"key":"e_1_3_3_108_2","first-page":"3918","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Oord Aaron","year":"2018","unstructured":"Aaron Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George Driessche, Edward Lockhart, Luis Cobo, Florian Stimberg, et\u00a0al. 2018. Parallel WaveNet: Fast high-fidelity speech synthesis. In Proceedings of the International Conference on Machine Learning. PMLR, 3918\u20133926."},{"key":"e_1_3_3_109_2","article-title":"Im2Text: Describing images using 1 million captioned photographs","volume":"24","author":"Ordonez Vicente","year":"2011","unstructured":"Vicente Ordonez, Girish Kulkarni, and Tamara Berg. 2011. Im2Text: Describing images using 1 million captioned photographs. Adv. Neural Inf. Process. Syst. 24 (2011).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_3_110_2","first-page":"5206","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Panayotov Vassil","year":"2015","unstructured":"Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. LibriSpeech: An ASR corpus based on public domain audio books. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 5206\u20135210."},{"key":"e_1_3_3_111_2","doi-asserted-by":"crossref","first-page":"511","DOI":"10.1007\/978-0-85729-997-0_26","volume-title":"Visual Analysis of Humans","author":"Pantic Maja","year":"2011","unstructured":"Maja Pantic, Roderick Cowie, Francesca D\u2019Errico, Dirk Heylen, Marc Mehu, Catherine Pelachaud, Isabella Poggi, Marc Schroeder, and Alessandro Vinciarelli. 2011. Social signal processing: The research agenda. In Visual Analysis of Humans. Springer, 511\u2013538."},{"key":"e_1_3_3_112_2","article-title":"Attentive explanations: Justifying decisions and pointing to the evidence","author":"Park Dong Huk","year":"2016","unstructured":"Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. 2016. Attentive explanations: Justifying decisions and pointing to the evidence. arXiv preprint arXiv:1612.04757 (2016).","journal-title":"arXiv preprint arXiv:1612.04757"},{"key":"e_1_3_3_113_2","first-page":"1577","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Patro Badri","year":"2020","unstructured":"Badri Patro, Shivansh Patel, and Vinay Namboodiri. 2020. Robust explanations for visual question answering. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 1577\u20131586."},{"key":"e_1_3_3_114_2","first-page":"8347","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Pei Wenjie","year":"2019","unstructured":"Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai. 2019. Memory-attended recurrent network for video captioning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 8347\u20138356."},{"key":"e_1_3_3_115_2","first-page":"3039","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Perez-Martin Jesus","year":"2021","unstructured":"Jesus Perez-Martin, Benjamin Bustos, and Jorge P\u00e9rez. 2021. Improving video captioning with temporal composition of a visual-syntactic embedding. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 3039\u20133049."},{"key":"e_1_3_3_116_2","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1007\/978-1-4613-0403-6_33","volume-title":"Multimedia Communications and Video Coding","author":"Petajan Eric","year":"1996","unstructured":"Eric Petajan and Hans Peter Graf. 1996. Automatic lipreading research: Historic overview and current work. In Multimedia Communications and Video Coding. Springer, 265\u2013275."},{"key":"e_1_3_3_117_2","article-title":"Deep voice 3: Scaling text-to-speech with convolutional sequence learning","author":"Ping Wei","year":"2017","unstructured":"Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller. 2017. Deep voice 3: Scaling text-to-speech with convolutional sequence learning. arXiv preprint arXiv:1710.07654 (2017).","journal-title":"arXiv preprint arXiv:1710.07654"},{"key":"e_1_3_3_118_2","article-title":"Semantically sensible video captioning (SSVC)","author":"Rahman Md","year":"2020","unstructured":"Md Rahman, Thasin Abedin, Khondokar S. S. Prottoy, Ayana Moshruba, Fazlul Hasan Siddiqui, et\u00a0al. 2020. Semantically sensible video captioning (SSVC). arXiv preprint arXiv:2009.07335 (2020).","journal-title":"arXiv preprint arXiv:2009.07335"},{"issue":"6","key":"e_1_3_3_119_2","doi-asserted-by":"crossref","first-page":"96","DOI":"10.1109\/MSP.2017.2738401","article-title":"Deep multimodal learning: A survey on recent advances and trends","volume":"34","author":"Ramachandram Dhanesh","year":"2017","unstructured":"Dhanesh Ramachandram and Graham W. Taylor. 2017. Deep multimodal learning: A survey on recent advances and trends. IEEE Sig. Process. Mag. 34, 6 (2017), 96\u2013108.","journal-title":"IEEE Sig. Process. Mag."},{"key":"e_1_3_3_120_2","first-page":"139","volume-title":"Proceedings of the NAACL HLT Workshop on Creating Speech and Language Data with Amazon\u2019s Mechanical Turk","author":"Rashtchian Cyrus","year":"2010","unstructured":"Cyrus Rashtchian, Peter Young, Micah Hodosh, and Julia Hockenmaier. 2010. Collecting image annotations using Amazon\u2019s Mechanical Turk. In Proceedings of the NAACL HLT Workshop on Creating Speech and Language Data with Amazon\u2019s Mechanical Turk. 139\u2013147."},{"key":"e_1_3_3_121_2","first-page":"25","article-title":"Grounding action descriptions in videos","volume":"1","author":"Regneri Michaela","year":"2013","unstructured":"Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal. 2013. Grounding action descriptions in videos. Trans. Assoc. Computat. Ling. 1 (2013), 25\u201336.","journal-title":"Trans. Assoc. Computat. Ling."},{"key":"e_1_3_3_122_2","article-title":"Exploring models and data for image question answering","volume":"28","author":"Ren Mengye","year":"2015","unstructured":"Mengye Ren, Ryan Kiros, and Richard Zemel. 2015. Exploring models and data for image question answering. Adv. Neural Inf. Process. Syst. 28 (2015).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_3_123_2","first-page":"1","volume-title":"Proceedings of the 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG)","author":"Ringeval Fabien","year":"2013","unstructured":"Fabien Ringeval, Andreas Sonderegger, Juergen Sauer, and Denis Lalanne. 2013. Introducing the RECOLA multimodal corpus of remote collaborative and affective interactions. In Proceedings of the 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG). IEEE, 1\u20138."},{"key":"e_1_3_3_124_2","doi-asserted-by":"crossref","first-page":"184","DOI":"10.1007\/978-3-319-11752-2_15","volume-title":"Proceedings of the German Conference on Pattern Recognition","author":"Rohrbach Anna","year":"2014","unstructured":"Anna Rohrbach, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and Bernt Schiele. 2014. Coherent multi-sentence video description with variable level of detail. In Proceedings of the German Conference on Pattern Recognition. Springer, 184\u2013195."},{"key":"e_1_3_3_125_2","doi-asserted-by":"crossref","first-page":"105596","DOI":"10.1016\/j.knosys.2020.105596","article-title":"A review of deep learning with special emphasis on architectures, applications and recent trends","volume":"194","author":"Sengupta Saptarshi","year":"2020","unstructured":"Saptarshi Sengupta, Sanchita Basak, Pallabi Saikia, Sayak Paul, Vasilios Tsalavoutis, Frederick Atiah, Vadlamani Ravi, and Alan Peters. 2020. A review of deep learning with special emphasis on architectures, applications and recent trends. Knowl.-Based syst. 194 (2020), 105596.","journal-title":"Knowl.-Based syst."},{"key":"e_1_3_3_126_2","first-page":"4779","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Shen Jonathan","year":"2018","unstructured":"Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, R. J. Skerrv-Ryan, et\u00a0al. 2018. Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 4779\u20134783."},{"key":"e_1_3_3_127_2","first-page":"510","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Sigurdsson Gunnar A.","year":"2016","unstructured":"Gunnar A. Sigurdsson, G\u00fcl Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016. Hollywood in homes: Crowdsourcing data collection for activity understanding. In Proceedings of the European Conference on Computer Vision. Springer, 510\u2013526."},{"issue":"1","key":"e_1_3_3_128_2","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/B:MTAP.0000046380.27575.a5","article-title":"Multimodal video indexing: A review of the state-of-the-art","volume":"25","author":"Snoek Cees G. M.","year":"2005","unstructured":"Cees G. M. Snoek and Marcel Worring. 2005. Multimodal video indexing: A review of the state-of-the-art. Multim. Tools Applic. 25, 1 (2005), 5\u201335.","journal-title":"Multim. Tools Applic."},{"key":"e_1_3_3_129_2","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1016\/j.jpdc.2020.10.001","article-title":"Online multimedia retrieval on CPU\u2013GPU platforms with adaptive work partition","volume":"148","author":"Souza Rafael","year":"2021","unstructured":"Rafael Souza, Andr\u00e9 Fernandes, Thiago S. F. X. Teixeira, George Teodoro, and Renato Ferreira. 2021. Online multimedia retrieval on CPU\u2013GPU platforms with adaptive work partition. J. Parallel Distrib. Comput. 148 (2021), 31\u201345.","journal-title":"J. Parallel Distrib. Comput."},{"key":"e_1_3_3_130_2","article-title":"VoiceLoop: Voice fitting and synthesis via a phonological loop","author":"Taigman Yaniv","year":"2018","unstructured":"Yaniv Taigman, Lior Wolf, Adam Polyak, and Eliya Nachmani. 2018. VoiceLoop: Voice fitting and synthesis via a phonological loop. arXiv preprint arXiv:1707.06588 (2018).","journal-title":"arXiv preprint arXiv:1707.06588"},{"key":"e_1_3_3_131_2","doi-asserted-by":"crossref","first-page":"523","DOI":"10.1145\/2556195.2556245","volume-title":"Proceedings of the 7th ACM International Conference on Web Search and Data Mining","author":"Tandon Niket","year":"2014","unstructured":"Niket Tandon, Gerard De Melo, Fabian Suchanek, and Gerhard Weikum. 2014. WebChild: Harvesting and organizing commonsense knowledge from the web. In Proceedings of the 7th ACM International Conference on Web Search and Data Mining. 523\u2013532."},{"key":"e_1_3_3_132_2","first-page":"1","article-title":"End-to-end audiovisual speech recognition system with multitask learning","volume":"23","author":"Tao Fei","year":"2020","unstructured":"Fei Tao and Carlos Busso. 2020. End-to-end audiovisual speech recognition system with multitask learning. IEEE Trans. Multim. 23 (2020), 1\u201311.","journal-title":"IEEE Trans. Multim."},{"key":"e_1_3_3_133_2","article-title":"Using descriptive video services to create a large data source for video annotation research","author":"Torabi Atousa","year":"2015","unstructured":"Atousa Torabi, Christopher Pal, Hugo Larochelle, and Aaron Courville. 2015. Using descriptive video services to create a large data source for video annotation research. arXiv preprint arXiv:1503.01070 (2015).","journal-title":"arXiv preprint arXiv:1503.01070"},{"key":"e_1_3_3_134_2","article-title":"Multi-modal emotion recognition on IEMOCAP with neural networks","author":"Tripathi Samarth","year":"2018","unstructured":"Samarth Tripathi and Homayoon Beigi. 2018. Multi-modal emotion recognition on IEMOCAP with neural networks. arXiv preprint arXiv:1804.05788 (2018).","journal-title":"arXiv preprint arXiv:1804.05788"},{"key":"e_1_3_3_135_2","first-page":"69","volume-title":"Proceedings of the IEEE Spoken Language Technology Workshop","author":"Tur Gokhan","year":"2008","unstructured":"Gokhan Tur, Andreas Stolcke, Lynn Voss, John Dowding, Beno\u00eet Favre, Raquel Fern\u00e1ndez, Matthew Frampton, Michael Frandsen, Clint Frederickson, Martin Graciarena, et\u00a0al. 2008. The CALO meeting speech recognition and understanding system. In Proceedings of the IEEE Spoken Language Technology Workshop. IEEE, 69\u201372."},{"key":"e_1_3_3_136_2","unstructured":"Christophe Veaux Junichi Yamagishi Kirsten MacDonald et\u00a0al. 2016. SUPERSEDED-CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit. (2016)."},{"key":"e_1_3_3_137_2","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/978-1-84882-054-8_1","volume-title":"Computers in the Human Interaction Loop","author":"Waibel Alex","year":"2009","unstructured":"Alex Waibel, Hartwig Steusloff, Rainer Stiefelhagen, and Kym Watson. 2009. Computers in the human interaction loop. In Computers in the Human Interaction Loop. Springer, 3\u20136."},{"key":"e_1_3_3_138_2","first-page":"496","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Wan Chia-Hung","year":"2019","unstructured":"Chia-Hung Wan, Shun-Po Chuang, and Hung-Yi Lee. 2019. Towards audio to scene image synthesis using generative adversarial network. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 496\u2013500."},{"key":"e_1_3_3_139_2","first-page":"7622","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Bairui","year":"2018","unstructured":"Bairui Wang, Lin Ma, Wei Zhang, and Wei Liu. 2018. Reconstruction network for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 7622\u20137631."},{"key":"e_1_3_3_140_2","doi-asserted-by":"crossref","first-page":"104543","DOI":"10.1109\/ACCESS.2020.2999568","article-title":"Cross-lingual image caption generation based on visual attention model","volume":"8","author":"Wang Bin","year":"2020","unstructured":"Bin Wang, Cungang Wang, Qian Zhang, Ying Su, Yang Wang, and Yanyan Xu. 2020. Cross-lingual image caption generation based on visual attention model. IEEE Access 8 (2020), 104543\u2013104554.","journal-title":"IEEE Access"},{"issue":"10","key":"e_1_3_3_141_2","doi-asserted-by":"crossref","first-page":"2413","DOI":"10.1109\/TPAMI.2017.2754246","article-title":"FVQA: Fact-based visual question answering","volume":"40","author":"Wang Peng","year":"2017","unstructured":"Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. 2017. FVQA: Fact-based visual question answering. IEEE Trans. Pattern Anal. Mach. Intell. 40, 10 (2017), 2413\u20132427.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_3_142_2","first-page":"1173","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Peng","year":"2017","unstructured":"Peng Wang, Qi Wu, Chunhua Shen, and Anton van den Hengel. 2017. The VQA-machine: Learning how to use existing vision algorithms to answer new questions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 1173\u20131182."},{"key":"e_1_3_3_143_2","first-page":"3081","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Wang Wei","year":"2018","unstructured":"Wei Wang, Yuxuan Ding, and Chunna Tian. 2018. A novel semantic attribute-based feature for image caption generation. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 3081\u20133085."},{"key":"e_1_3_3_144_2","first-page":"4213","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Wang Xin","year":"2018","unstructured":"Xin Wang, Wenhu Chen, Jiawei Wu, Yuan-Fang Wang, and William Yang Wang. 2018. Video captioning via hierarchical reinforcement learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4213\u20134222."},{"key":"e_1_3_3_145_2","doi-asserted-by":"crossref","first-page":"298","DOI":"10.1016\/j.ins.2020.08.009","article-title":"DRSL: Deep relational similarity learning for cross-modal retrieval","volume":"546","author":"Wang Xu","year":"2021","unstructured":"Xu Wang, Peng Hu, Liangli Zhen, and Dezhong Peng. 2021. DRSL: Deep relational similarity learning for cross-modal retrieval. Inf. Sci. 546 (2021), 298\u2013311.","journal-title":"Inf. Sci."},{"key":"e_1_3_3_146_2","article-title":"Tacotron: Towards end-to-end speech synthesis","author":"Wang Yuxuan","year":"2017","unstructured":"Yuxuan Wang, R. J. Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et\u00a0al. 2017. Tacotron: Towards end-to-end speech synthesis. arXiv preprint arXiv:1703.10135 (2017).","journal-title":"arXiv preprint arXiv:1703.10135"},{"key":"e_1_3_3_147_2","doi-asserted-by":"crossref","first-page":"102751","DOI":"10.1016\/j.jvcir.2020.102751","article-title":"Exploiting the local temporal information for video captioning","volume":"67","author":"Wei Ran","year":"2020","unstructured":"Ran Wei, Li Mi, Yaosi Hu, and Zhenzhong Chen. 2020. Exploiting the local temporal information for video captioning. J. Vis. Commun. Image Represent. 67 (2020), 102751.","journal-title":"J. Vis. Commun. Image Represent."},{"key":"e_1_3_3_148_2","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1016\/j.neucom.2019.12.073","article-title":"Multi-attention generative adversarial network for image captioning","volume":"387","author":"Wei Yiwei","year":"2020","unstructured":"Yiwei Wei, Leiquan Wang, Haiwen Cao, Mingwen Shao, and Chunlei Wu. 2020. Multi-attention generative adversarial network for image captioning. Neurocomputing 387 (2020), 91\u201399.","journal-title":"Neurocomputing"},{"issue":"3","key":"e_1_3_3_149_2","first-page":"1250","article-title":"Spatiotemporal multimodal learning with 3D CNNs for video action recognition","volume":"32","author":"Wu Hanbo","year":"2021","unstructured":"Hanbo Wu, Xin Ma, and Yibin Li. 2021. Spatiotemporal multimodal learning with 3D CNNs for video action recognition. IEEE Trans. Circ. Syst. Vid. Technol. 32, 3 (2021), 1250\u20131261.","journal-title":"IEEE Trans. Circ. Syst. Vid. Technol."},{"issue":"25","key":"e_1_3_3_150_2","doi-asserted-by":"crossref","first-page":"1642","DOI":"10.1049\/el.2017.3159","article-title":"Cascade recurrent neural network for image caption generation","volume":"53","author":"Wu Jie","year":"2017","unstructured":"Jie Wu and Haifeng Hu. 2017. Cascade recurrent neural network for image caption generation. Electron. Lett. 53, 25 (2017), 1642\u20131643.","journal-title":"Electron. Lett."},{"key":"e_1_3_3_151_2","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1016\/j.cviu.2017.05.001","article-title":"Visual question answering: A survey of methods and datasets","volume":"163","author":"Wu Qi","year":"2017","unstructured":"Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. 2017. Visual question answering: A survey of methods and datasets. Comput. Vis. Image Underst. 163 (2017), 21\u201340.","journal-title":"Comput. Vis. Image Underst."},{"key":"e_1_3_3_152_2","first-page":"115648","article-title":"Visual question answering model based on visual relationship detection","volume":"80","author":"Xi Yuling","year":"2020","unstructured":"Yuling Xi, Yanning Zhang, Songtao Ding, and Shaohua Wan. 2020. Visual question answering model based on visual relationship detection. Sig. Process.: Image Commun. 80 (2020), 115648.","journal-title":"Sig. Process.: Image Commun."},{"key":"e_1_3_3_153_2","first-page":"5288","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Xu Jun","year":"2016","unstructured":"Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. MSR-VTT: A large video description dataset for bridging video and language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5288\u20135296."},{"key":"e_1_3_3_154_2","first-page":"1772","article-title":"Deep reinforcement polishing network for video captioning","volume":"23","author":"Xu Wanru","year":"2020","unstructured":"Wanru Xu, Jian Yu, Zhenjiang Miao, Lili Wan, Yi Tian, and Qiang Ji. 2020. Deep reinforcement polishing network for video captioning. IEEE Trans. Multim. 23 (2020), 1772\u20131784.","journal-title":"IEEE Trans. Multim."},{"issue":"5","key":"e_1_3_3_155_2","first-page":"1243","article-title":"Shared multi-view data representation for multi-domain event detection","volume":"42","author":"Yang Zhenguo","year":"2019","unstructured":"Zhenguo Yang, Qing Li, Wenyin Liu, and Jianming Lv. 2019. Shared multi-view data representation for multi-domain event detection. IEEE Trans. Pattern Anal. Mach. Intell. 42, 5 (2019), 1243\u20131256.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"4","key":"e_1_3_3_156_2","doi-asserted-by":"crossref","first-page":"e0226248","DOI":"10.1371\/journal.pone.0226248","article-title":"Multimodal mental health analysis in social media","volume":"15","author":"Yazdavar Amir Hossein","year":"2020","unstructured":"Amir Hossein Yazdavar, Mohammad Saeid Mahdavinejad, Goonmeet Bajaj, William Romine, Amit Sheth, Amir Hassan Monadjemi, Krishnaprasad Thirunarayan, John M. Meddar, Annie Myers, Jyotishman Pathak, et\u00a0al. 2020. Multimodal mental health analysis in social media. PLoS One 15, 4 (2020), e0226248.","journal-title":"PLoS One"},{"key":"e_1_3_3_157_2","first-page":"67","article-title":"From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions","volume":"2","author":"Young Peter","year":"2014","unstructured":"Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Trans. Assoc. Computat. Ling. 2 (2014), 67\u201378.","journal-title":"Trans. Assoc. Computat. Ling."},{"key":"e_1_3_3_158_2","doi-asserted-by":"crossref","first-page":"107563","DOI":"10.1016\/j.patcog.2020.107563","article-title":"Cross-modal knowledge reasoning for knowledge-based visual question answering","volume":"108","author":"Yu Jing","year":"2020","unstructured":"Jing Yu, Zihao Zhu, Yujing Wang, Weifeng Zhang, Yue Hu, and Jianlong Tan. 2020. Cross-modal knowledge reasoning for knowledge-based visual question answering. Pattern Recog. 108 (2020), 107563.","journal-title":"Pattern Recog."},{"key":"e_1_3_3_159_2","first-page":"6281","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yu Zhou","year":"2019","unstructured":"Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. 2019. Deep modular co-attention networks for visual question answering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6281\u20136290."},{"key":"e_1_3_3_160_2","first-page":"1821","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Yu Zhou","year":"2017","unstructured":"Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao. 2017. Multi-modal factorized bilinear pooling with co-attention learning for visual question answering. In Proceedings of the IEEE International Conference on Computer Vision. 1821\u20131830."},{"key":"e_1_3_3_161_2","first-page":"115731","article-title":"Correlation Net: Spatiotemporal multimodal deep learning for action recognition","volume":"82","author":"Yudistira Novanto","year":"2020","unstructured":"Novanto Yudistira and Takio Kurita. 2020. Correlation Net: Spatiotemporal multimodal deep learning for action recognition. Sig. Process.: Image Commun. 82 (2020), 115731.","journal-title":"Sig. Process.: Image Commun."},{"issue":"11","key":"e_1_3_3_162_2","doi-asserted-by":"crossref","first-page":"65","DOI":"10.1109\/35.41402","article-title":"Integration of acoustic and visual speech signals using neural networks","volume":"27","author":"Yuhas Ben P.","year":"1989","unstructured":"Ben P. Yuhas, Moise H. Goldstein, and Terrence J. Sejnowski. 1989. Integration of acoustic and visual speech signals using neural networks. IEEE Commun. Mag. 27, 11 (1989), 65\u201371.","journal-title":"IEEE Commun. Mag."},{"issue":"3","key":"e_1_3_3_163_2","doi-asserted-by":"crossref","first-page":"478","DOI":"10.1109\/JSTSP.2020.2987728","article-title":"Multimodal intelligence: Representation learning, information fusion, and applications","volume":"14","author":"Zhang Chao","year":"2020","unstructured":"Chao Zhang, Zichao Yang, Xiaodong He, and Li Deng. 2020. Multimodal intelligence: Representation learning, information fusion, and applications. IEEE J. Select. Topics Sig. Process. 14, 3 (2020), 478\u2013493.","journal-title":"IEEE J. Select. Topics Sig. Process."},{"key":"e_1_3_3_164_2","first-page":"1","volume-title":"Proceedings of the International Conference on Machine Learning and Cybernetics (ICMLC)","author":"Zhang Su-Fang","year":"2019","unstructured":"Su-Fang Zhang, Jun-Hai Zhai, Bo-Jun Xie, Yan Zhan, and Xin Wang. 2019. Multimodal representation learning: Advances, trends and challenges. In Proceedings of the International Conference on Machine Learning and Cybernetics (ICMLC). IEEE, 1\u20136."},{"issue":"12","key":"e_1_3_3_165_2","doi-asserted-by":"crossref","first-page":"3088","DOI":"10.1109\/TPAMI.2019.2920899","article-title":"Reconstruct and represent video contents for captioning via reinforcement learning","volume":"42","author":"Zhang Wei","year":"2019","unstructured":"Wei Zhang, Bairui Wang, Lin Ma, and Wei Liu. 2019. Reconstruct and represent video contents for captioning via reinforcement learning. IEEE Trans. Pattern Anal. Mach. Intell. 42, 12 (2019), 3088\u20133101.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"7","key":"e_1_3_3_166_2","doi-asserted-by":"crossref","first-page":"1681","DOI":"10.1109\/TMM.2018.2888822","article-title":"High-quality image captioning with fine-grained and semantic-guided visual attention","volume":"21","author":"Zhang Zongjian","year":"2018","unstructured":"Zongjian Zhang, Qiang Wu, Yang Wang, and Fang Chen. 2018. High-quality image captioning with fine-grained and semantic-guided visual attention. IEEE Trans. Multim. 21, 7 (2018), 1681\u20131693.","journal-title":"IEEE Trans. Multim."},{"key":"e_1_3_3_167_2","doi-asserted-by":"crossref","first-page":"1799","DOI":"10.1109\/TMM.2020.3003592","article-title":"Dense video captioning using graph-based sentence summarization","volume":"23","author":"Zhang Zhiwang","year":"2020","unstructured":"Zhiwang Zhang, Dong Xu, Wanli Ouyang, and Luping Zhou. 2020. Dense video captioning using graph-based sentence summarization. IEEE Trans. Multim. 23 (2020), 1799\u20131810.","journal-title":"IEEE Trans. Multim."},{"key":"e_1_3_3_168_2","first-page":"10394","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhen Liangli","year":"2019","unstructured":"Liangli Zhen, Peng Hu, Xu Wang, and Dezhong Peng. 2019. Deep supervised cross-modal retrieval. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10394\u201310403."},{"key":"e_1_3_3_169_2","first-page":"9299","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Zhou Hang","year":"2019","unstructured":"Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang. 2019. Talking face generation by adversarially disentangled audio-visual representation. In Proceedings of the AAAI Conference on Artificial Intelligence. 9299\u20139306."},{"key":"e_1_3_3_170_2","first-page":"4176","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhou Hang","year":"2021","unstructured":"Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu. 2021. Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 4176\u20134186."},{"key":"e_1_3_3_171_2","first-page":"3550","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhou Yipin","year":"2018","unstructured":"Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L. Berg. 2018. Visual to sound: Generating natural sound for videos in the wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3550\u20133558."},{"issue":"3","key":"e_1_3_3_172_2","first-page":"1","article-title":"Physiological signals-based emotion recognition via high-order correlation learning","volume":"15","author":"Zhu Junjie","year":"2019","unstructured":"Junjie Zhu, Yuxuan Wei, Yifan Feng, Xibin Zhao, and Yue Gao. 2019. Physiological signals-based emotion recognition via high-order correlation learning. ACM Trans. Multim. Comput. Commun. Applic. 15, 3s (2019), 1\u201318.","journal-title":"ACM Trans. Multim. Comput. Commun. Applic."},{"key":"e_1_3_3_173_2","first-page":"4995","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhu Yuke","year":"2016","unstructured":"Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2016. Visual7W: Grounded question answering in images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4995\u20135004."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3545572","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3545572","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:02:45Z","timestamp":1750186965000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3545572"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,17]]},"references-count":172,"journal-issue":{"issue":"2s","published-print":{"date-parts":[[2023,6,30]]}},"alternative-id":["10.1145\/3545572"],"URL":"https:\/\/doi.org\/10.1145\/3545572","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,17]]},"assertion":[{"value":"2021-12-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-05-31","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}