{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,4]],"date-time":"2026-08-04T04:11:30Z","timestamp":1785816690682,"version":"3.56.0"},"reference-count":217,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,10,23]],"date-time":"2023-10-23T00:00:00Z","timestamp":1698019200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,3,31]]},"abstract":"<jats:p>\n            Multimodality Representation Learning, as a technique of learning to embed information from different modalities and their correlations, has achieved remarkable success on a variety of applications, such as Visual Question Answering (VQA), Natural Language for Visual Reasoning (NLVR), and Vision Language Retrieval (VLR). Among these applications, cross-modal interaction and complementary information from different modalities are crucial for advanced models to perform any multimodal task, e.g., understand, recognize, retrieve, or generate optimally. Researchers have proposed diverse methods to address these tasks. The different variants of transformer-based architectures performed extraordinarily on multiple modalities. This survey presents the comprehensive literature on the evolution and enhancement of deep learning multimodal architectures to deal with textual, visual and audio features for diverse cross-modal and modern multimodal tasks. This study summarizes the (\n            <jats:italic>i<\/jats:italic>\n            )\u00a0recent task-specific deep learning methodologies, (\n            <jats:italic>ii<\/jats:italic>\n            )\u00a0the pretraining types and multimodal pretraining objectives, (\n            <jats:italic>iii<\/jats:italic>\n            )\u00a0from state-of-the-art pretrained multimodal approaches to unifying architectures, and (\n            <jats:italic>iv<\/jats:italic>\n            )\u00a0multimodal task categories and possible future improvements that can be devised for better multimodal learning. Moreover, we prepare a dataset section for new researchers that covers most of the benchmarks for pretraining and finetuning. Finally, major challenges, gaps, and potential research topics are explored. A constantly-updated paperlist related to our survey is maintained at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/marslanm\/multimodality-representation-learning\">https:\/\/github.com\/marslanm\/multimodality-representation-learning<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3617833","type":"journal-article","created":{"date-parts":[[2023,8,29]],"date-time":"2023-08-29T11:24:17Z","timestamp":1693308257000},"page":"1-34","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":52,"title":["Multimodality Representation Learning: A Survey on Evolution, Pretraining and Its Applications"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8773-9972","authenticated-orcid":false,"given":"Muhammad Arslan","family":"Manzoor","sequence":"first","affiliation":[{"name":"Mohamed bin Zayed University of Artificial Intelligence, UAE"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3203-4307","authenticated-orcid":false,"given":"Sarah","family":"Albarri","sequence":"additional","affiliation":[{"name":"Mohamed bin Zayed University of Artificial Intelligence, UAE"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4116-6623","authenticated-orcid":false,"given":"Ziting","family":"Xian","sequence":"additional","affiliation":[{"name":"Sun Yat-sen University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5374-0318","authenticated-orcid":false,"given":"Zaiqiao","family":"Meng","sequence":"additional","affiliation":[{"name":"University of Glasgow, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3600-1510","authenticated-orcid":false,"given":"Preslav","family":"Nakov","sequence":"additional","affiliation":[{"name":"Mohamed bin Zayed University of Artificial Intelligence, UAE"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1625-2168","authenticated-orcid":false,"given":"Shangsong","family":"Liang","sequence":"additional","affiliation":[{"name":"Mohamed bin Zayed University of Artificial Intelligence, UAE"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,10,23]]},"reference":[{"key":"e_1_3_1_2_2","article-title":"Recent advances and trends in multimodal deep learning: A review","author":"Summaira Jabeen","year":"2021","unstructured":"Jabeen Summaira, Xi Li, Amin Muhammad Shoib, Songyuan Li, and Jabbar Abdul. 2021. Recent advances and trends in multimodal deep learning: A review. arXiv preprint arXiv:2105.11087 (2021).","journal-title":"arXiv preprint arXiv:2105.11087"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2798607"},{"key":"e_1_3_1_4_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Shi Bowen","year":"2022","unstructured":"Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrahman Mohamed. 2022. Learning audio-visual speech representation by masked multimodal cluster prediction. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.516"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2916887"},{"key":"e_1_3_1_8_2","first-page":"1","article-title":"A survey on deep multimodal learning for computer vision: Advances, trends, applications, and datasets","author":"Bayoudh Khaled","year":"2021","unstructured":"Khaled Bayoudh, Raja Knani, Fay\u00e7al Hamdaoui, and Abdellatif Mtibaa. 2021. A survey on deep multimodal learning for computer vision: Advances, trends, applications, and datasets. The Visual Computer (2021), 1\u201332.","journal-title":"The Visual Computer"},{"key":"e_1_3_1_9_2","article-title":"VisualBERT: A simple and performant baseline for vision and language","author":"Li Liunian Harold","year":"2019","unstructured":"Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. VisualBERT: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557 (2019).","journal-title":"arXiv preprint arXiv:1908.03557"},{"key":"e_1_3_1_10_2","article-title":"ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks","volume":"32","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01045"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1038\/264746a0"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-010-0182-0"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2013.2267205"},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","first-page":"290","DOI":"10.1117\/12.333848","volume-title":"Storage and Retrieval for Image and Video Databases VII","author":"Lienhart Rainer W.","year":"1998","unstructured":"Rainer W. Lienhart. 1998. Comparison of automatic shot boundary detection algorithms. In Storage and Retrieval for Image and Video Databases VII, Vol. 3656. SPIE, 290\u2013301."},{"key":"e_1_3_1_16_2","doi-asserted-by":"crossref","unstructured":"Jean Carletta Simone Ashby Sebastien Bourban Mike Flynn Mael Guillemot Thomas Hain Jaroslav Kadlec Vasilis Karaiskos Wessel Kraaij Melissa Kronenthal Guillaume Lathoud Mike Lincoln Masson Agnes Lisowska Iain McCowan Wilfried Post Dennis Reidsma and Pierre D. Wellner. 2006. The AMI meeting corpus: A pre-announcement. In Machine Learning for Multimodal Interaction: Second International Workshop (MLMI 2005 Edinburgh UK July 11-13 2005 Revised Selected Papers 2) Springer 28\u201339.","DOI":"10.1007\/11677482_3"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2010.5583006"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-24571-8_53"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/2661806.2661807"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.168"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1873987"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2627563"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_3_1_24_2","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. International Conference on Learning Representations (2015).","journal-title":"International Conference on Learning Representations"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_28_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","unstructured":"Zhe Gan Linjie Li Chunyuan Li Lijuan Wang Zicheng Liu and Jianfeng Gao. 2022. Vision-language pre-training: Basics recent advances and future trends. Foundations and Trends\u00ae in Computer Graphics and Vision 14 3-4 (2022) 163\u2013352.","DOI":"10.1561\/0600000105"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_44"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00637"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.01039"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00680"},{"key":"e_1_3_1_34_2","first-page":"32897","article-title":"VLMo: Unified vision-language pre-training with mixture-of-modality-experts","volume":"35","author":"Bao Hangbo","year":"2022","unstructured":"Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, and Furu Wei. 2022. VLMo: Unified vision-language pre-training with mixture-of-modality-experts. Advances in Neural Information Processing Systems 35 (2022), 32897\u201332912.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_35_2","first-page":"5583","volume-title":"International Conference on Machine Learning","author":"Kim Wonjae","year":"2021","unstructured":"Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. ViLT: Vision-and-language transformer without convolution or region supervision. In International Conference on Machine Learning. PMLR, 5583\u20135594."},{"key":"e_1_3_1_36_2","first-page":"4171","volume-title":"Proceedings of the NAACL-HLT","author":"Kenton Jacob Devlin Ming-Wei Chang","year":"2019","unstructured":"Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the NAACL-HLT. 4171\u20134186."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"e_1_3_1_38_2","first-page":"12449","article-title":"wav2vec 2.0: A framework for self-supervised learning of speech representations","volume":"33","author":"Baevski Alexei","year":"2020","unstructured":"Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems 33 (2020), 12449\u201312460.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2021.12.003"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco_a_01273"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11633-022-1369-5"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2022\/762"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2022\/773"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2017.2738401"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1.11688"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2017.05.001"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2021.103982"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.3390\/app12042204"},{"key":"e_1_3_1_49_2","article-title":"The multimodal sentiment analysis in car reviews (muse-car) dataset: Collection, insights and improvements","author":"Stappen Lukas","year":"2021","unstructured":"Lukas Stappen, Alice Baird, Lea Schumann, and Schuller Bjorn. 2021. The multimodal sentiment analysis in car reviews (muse-car) dataset: Collection, insights and improvements. IEEE Transactions on Affective Computing (2021).","journal-title":"IEEE Transactions on Affective Computing"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1002\/widm.1415"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3042066"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1049\/el.2017.3159"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i2.16258"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_1_55_2","first-page":"689","volume-title":"Proceedings of the 28th International Conference on Machine Learning (ICML\u201911)","author":"Ngiam Jiquan","year":"2011","unstructured":"Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y. Ng. 2011. Multimodal deep learning. In Proceedings of the 28th International Conference on Machine Learning (ICML\u201911). 689\u2013696."},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390294"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2022.3224228"},{"key":"e_1_3_1_58_2","volume-title":"Proceedings of International Conference on Learning Representations","author":"Li Yujia","year":"2016","unstructured":"Yujia Li, Richard Zemel, Marc Brockschmidt, and Daniel Tarlow. 2016. Gated graph sequence neural networks. In Proceedings of International Conference on Learning Representations."},{"key":"e_1_3_1_59_2","unstructured":"Petar Velickovic Guillem Cucurull Arantxa Casanova Adriana Romero Pietro Lio and Yoshua Bengio. 2017. Graph attention networks. Stat 1050 20 (2017) 10\u201348550."},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.273"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01276"},{"key":"e_1_3_1_62_2","article-title":"VQA-GNN: Reasoning with multimodal semantic graph for visual question answering","author":"Wang Yanan","year":"2022","unstructured":"Yanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada, and Jure Leskovec. 2022. VQA-GNN: Reasoning with multimodal semantic graph for visual question answering. arXiv preprint arXiv:2205.11501 (2022).","journal-title":"arXiv preprint arXiv:2205.11501"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3067607"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.657"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00443"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467327"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i5.20492"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00553"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467206"},{"key":"e_1_3_1_70_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Clark Kevin","year":"2020","unstructured":"Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: Pre-training text encoders as discriminators rather than generators. In Proceedings of the International Conference on Learning Representations. OpenReview.net."},{"key":"e_1_3_1_71_2","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692 (2019).","journal-title":"arXiv preprint arXiv:1907.11692"},{"key":"e_1_3_1_72_2","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btz682"},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.woah-1.3"},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.280"},{"key":"e_1_3_1_75_2","article-title":"Pre-training technique to localize medical BERT and enhance biomedical BERT","author":"Wada Shoya","year":"2020","unstructured":"Shoya Wada, Toshihiro Takeda, Shiro Manabe, Shozo Konishi, Jun Kamohara, and Yasushi Matsumura. 2020. Pre-training technique to localize medical BERT and enhance biomedical BERT. arXiv preprint arXiv:2005.07202 (2020).","journal-title":"arXiv preprint arXiv:2005.07202"},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.740"},{"key":"e_1_3_1_77_2","doi-asserted-by":"crossref","unstructured":"Yujia Qin Yankai Lin Jing Yi Jiajie Zhang Xu Han Zhengyan Zhang Yusheng Su Zhiyuan Liu Peng Li Maosong Sun and Jie Zhou. 2022. Knowledge Inheritance for Pre-trained Language Models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 3921\u20133937.","DOI":"10.18653\/v1\/2022.naacl-main.288"},{"key":"e_1_3_1_78_2","article-title":"Improving language understanding by generative pre-training","author":"Radford Alec","year":"2018","unstructured":"Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre-training. The University of British Columbia (2018).","journal-title":"The University of British Columbia"},{"key":"e_1_3_1_79_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-srw.12"},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01157"},{"key":"e_1_3_1_81_2","article-title":"InterBERT: Vision-and-language interaction for multi-modal pretraining","author":"Lin Junyang","year":"2020","unstructured":"Junyang Lin, An Yang, Yichang Zhang, Jie Liu, Jingren Zhou, and Hongxia Yang. 2020. InterBERT: Vision-and-language interaction for multi-modal pretraining. arXiv preprint arXiv:2003.13198 (2020).","journal-title":"arXiv preprint arXiv:2003.13198"},{"key":"e_1_3_1_82_2","first-page":"1589","volume-title":"Findings of the Association for Computational Linguistics: NAACL 2022","author":"Liu Yongfei","year":"2022","unstructured":"Yongfei Liu, Chenfei Wu, Shao-Yen Tseng, Vasudev Lal, Xuming He, and Nan Duan. 2022. KD-VLP: Improving end-to-end vision-and-language pretraining with object knowledge distillation. In Findings of the Association for Computational Linguistics: NAACL 2022. 1589\u20131600."},{"key":"e_1_3_1_83_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_7"},{"key":"e_1_3_1_84_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.161"},{"key":"e_1_3_1_85_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00414"},{"key":"e_1_3_1_86_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Lan Zhenzhong","year":"2019","unstructured":"Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. ALBERT: A lite BERT for self-supervised learning of language representations. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_87_2","doi-asserted-by":"publisher","DOI":"10.5555\/3455716.3455856"},{"key":"e_1_3_1_88_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_1_89_2","volume-title":"International Conference on Learning Representations","author":"Zhu Xizhou","year":"2021","unstructured":"Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2021. Deformable DETR: Deformable transformers for end-to-end object detection. In International Conference on Learning Representations."},{"key":"e_1_3_1_90_2","unstructured":"Alec Radford Jong Wook Kim Chris Hallacy Aditya Ramesh Gabriel Goh Sandhini Agarwal Girish Sastry Amanda Askell Pamela Mishkin Jack Clark Gretchen Krueger and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning PMLR 8748\u20138763."},{"key":"e_1_3_1_91_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1238"},{"key":"e_1_3_1_92_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.7005"},{"key":"e_1_3_1_93_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.303"},{"key":"e_1_3_1_94_2","doi-asserted-by":"crossref","unstructured":"Xiujun Li Xi Yin Chunyuan Li Pengchuan Zhang Xiaowei Hu Lei Zhang Lijuan Wang Houdong Hu Li Dong Furu Wei Yejin Choi and Jianfeng Gao. 2020. Oscar: Object-semantics aligned pre-training for vision-language tasks. In European Conference on Computer Vision Springer 121\u2013137.","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"e_1_3_1_95_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01101"},{"key":"e_1_3_1_96_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i4.16431"},{"key":"e_1_3_1_97_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.162"},{"key":"e_1_3_1_98_2","first-page":"12888","volume-title":"Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research)","volume":"162","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the 39th International Conference on Machine Learning (Proceedings of Machine Learning Research), Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato (Eds.), Vol. 162. PMLR, 12888\u201312900. https:\/\/proceedings.mlr.press\/v162\/li22n.html"},{"key":"e_1_3_1_99_2","article-title":"BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597 (2023).","journal-title":"arXiv preprint arXiv:2301.12597"},{"key":"e_1_3_1_100_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9414460"},{"key":"e_1_3_1_101_2","article-title":"Distributed representations of words and phrases and their compositionality","volume":"26","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. Advances in Neural Information Processing Systems 26 (2013).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_102_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-2501"},{"key":"e_1_3_1_103_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.463"},{"key":"e_1_3_1_104_2","doi-asserted-by":"crossref","unstructured":"Yonatan Bisk Ari Holtzman Jesse Thomason Jacob Andreas Yoshua Bengio Joyce Chai Mirella Lapata Angeliki Lazaridou Jonathan May Aleksandr Nisnevich Nicoloas Pinto and Joseph P. Turian. 2020. Experience Grounds Language. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201920) . 8718\u20138735.","DOI":"10.18653\/v1\/2020.emnlp-main.703"},{"key":"e_1_3_1_105_2","doi-asserted-by":"crossref","unstructured":"Ranjay Krishna Yuke Zhu Oliver Groth Justin Johnson Kenji Hata Joshua Kravitz Stephanie Chen Yannis Kalantidis Li-Jia Li David A. Shamma Michael S. Bernstein and Li Fei-Fei. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International Journal of Computer Vision 123 (2017) 32\u201373.","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_1_106_2","doi-asserted-by":"publisher","DOI":"10.1145\/2812802"},{"key":"e_1_3_1_107_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00904"},{"key":"e_1_3_1_108_2","article-title":"Towards a unified foundation model: Jointly pre-training transformers on unpaired images and text","author":"Li Qing","year":"2021","unstructured":"Qing Li, Boqing Gong, Yin Cui, Dan Kondratyuk, Xianzhi Du, Ming-Hsuan Yang, and Matthew Brown. 2021. Towards a unified foundation model: Jointly pre-training transformers on unpaired images and text. arXiv preprint arXiv:2112.07074 (2021).","journal-title":"arXiv preprint arXiv:2112.07074"},{"key":"e_1_3_1_109_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00147"},{"key":"e_1_3_1_110_2","article-title":"VATT: Transformers for multimodal self-supervised learning from raw video, audio and text","volume":"34","author":"Akbari Hassan","year":"2021","unstructured":"Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. 2021. VATT: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in Neural Information Processing Systems 34 (2021).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_111_2","first-page":"23318","volume-title":"International Conference on Machine Learning","author":"Wang Peng","year":"2022","unstructured":"Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. OFA: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning. PMLR, 23318\u201323340."},{"key":"e_1_3_1_112_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"e_1_3_1_113_2","article-title":"InstructBLIP: Towards general-purpose vision-language models with instruction tuning","author":"Dai Wenliang","year":"2023","unstructured":"Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. arXiv preprint arXiv:2305.06500 (2023).","journal-title":"arXiv preprint arXiv:2305.06500"},{"key":"e_1_3_1_114_2","unstructured":"Hyung Won Chung Le Hou Shayne Longpre Barret Zoph Yi Tay William Fedus Eric Li Xuezhi Wang Mostafa Dehghani Siddhartha Brahma Albert Webson Shixiang Shane Gu Zhuyun Dai Mirac Suzgun Xinyun Chen Aakanksha Chowdhery Dasha Valter Sharan Narang Gaurav Mishra Adams Wei Yu Vincent Zhao Yanping Huang Andrew M. Dai Hongkun Yu Slav Petrov Ed Huai-hsin Chi Jeff Dean Jacob Devlin Adam Roberts Denny Zhou Quoc V. Le and Jason Wei. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416 (2022)."},{"key":"e_1_3_1_115_2","unstructured":"Wei-Lin Chiang Zhuohan Li Zi Lin Ying Sheng Zhanghao Wu Hao Zhang Lianmin Zheng Siyuan Zhuang Yonghao Zhuang Joseph E. Gonzalez Ion Stoica and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. (March2023). https:\/\/lmsys.org\/blog\/2023-03-30-vicuna\/"},{"key":"e_1_3_1_116_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.33"},{"key":"e_1_3_1_117_2","unstructured":"Yang Xu Yiheng Xu Tengchao Lv Lei Cui Furu Wei Guoxin Wang Yijuan Lu Dinei Florencio Cha Zhang Wanxiang Che Min Zhang and Lidong Zhou. 2021. LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . 2579\u20132591."},{"key":"e_1_3_1_118_2","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475345"},{"key":"e_1_3_1_119_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDAR.2019.00244"},{"key":"e_1_3_1_120_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDARW.2019.10029"},{"key":"e_1_3_1_121_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00454"},{"issue":"8","key":"e_1_3_1_122_2","article-title":"Do we really need explicit position encodings for vision transformers","volume":"3","author":"Chu Xiangxiang","year":"2021","unstructured":"Xiangxiang Chu, Bo Zhang, Zhi Tian, Xiaolin Wei, and Huaxia Xia. 2021. Do we really need explicit position encodings for vision transformers. arXiv preprint arXiv:2102.10882 3, 8 (2021).","journal-title":"arXiv preprint arXiv:2102.10882"},{"key":"e_1_3_1_123_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.144"},{"key":"e_1_3_1_124_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1659"},{"key":"e_1_3_1_125_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.326"},{"key":"e_1_3_1_126_2","volume-title":"Conference on Neural Information Processing Systems","author":"Sanabria Ramon","year":"2018","unstructured":"Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Lo\u00efc Barrault, Lucia Specia, and Florian Metze. 2018. How2: A large-scale dataset for multimodal language understanding. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_127_2","article-title":"DeCoAR 2.0: Deep contextualized acoustic representations with vector quantization","author":"Ling Shaoshi","year":"2020","unstructured":"Shaoshi Ling and Yuzong Liu. 2020. DeCoAR 2.0: Deep contextualized acoustic representations with vector quantization. arXiv preprint arXiv:2012.06659 (2020).","journal-title":"arXiv preprint arXiv:2012.06659"},{"key":"e_1_3_1_128_2","article-title":"LRS3-TED: A large-scale dataset for visual speech recognition","author":"Afouras Triantafyllos","year":"2018","unstructured":"Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman. 2018. LRS3-TED: A large-scale dataset for visual speech recognition. arXiv preprint arXiv:1809.00496 (2018).","journal-title":"arXiv preprint arXiv:1809.00496"},{"key":"e_1_3_1_129_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU46091.2019.9004036"},{"key":"e_1_3_1_130_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01381"},{"key":"e_1_3_1_131_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2017.2778151"},{"key":"e_1_3_1_132_2","doi-asserted-by":"publisher","DOI":"10.1109\/34.927467"},{"key":"e_1_3_1_133_2","first-page":"1","volume-title":"2021 9th European Workshop on Visual Information Processing (EUVIP)","author":"Zhu Lingyu","year":"2021","unstructured":"Lingyu Zhu and Esa Rahtu. 2021. Leveraging category information for single-frame visual sound source separation. In 2021 9th European Workshop on Visual Information Processing (EUVIP). IEEE, 1\u20136."},{"key":"e_1_3_1_134_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_35"},{"key":"e_1_3_1_135_2","doi-asserted-by":"publisher","DOI":"10.1177\/0165551515608733"},{"key":"e_1_3_1_136_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-5406"},{"key":"e_1_3_1_137_2","doi-asserted-by":"publisher","DOI":"10.1186\/s12992-021-00667-7"},{"key":"e_1_3_1_138_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.coling-main.93"},{"key":"e_1_3_1_139_2","doi-asserted-by":"publisher","DOI":"10.5555\/3163580.3163638"},{"key":"e_1_3_1_140_2","doi-asserted-by":"publisher","DOI":"10.1007\/s42979-021-00971-4"},{"key":"e_1_3_1_141_2","doi-asserted-by":"publisher","DOI":"10.1609\/icwsm.v12i1.14983"},{"key":"e_1_3_1_142_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2021.114939"},{"key":"e_1_3_1_143_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2019.03.005"},{"key":"e_1_3_1_144_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1565"},{"key":"e_1_3_1_145_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123454"},{"key":"e_1_3_1_146_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-47436-2_27"},{"key":"e_1_3_1_147_2","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219903"},{"key":"e_1_3_1_148_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313552"},{"key":"e_1_3_1_149_2","doi-asserted-by":"publisher","DOI":"10.3390\/app12031093"},{"key":"e_1_3_1_150_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00688"},{"key":"e_1_3_1_151_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2021.107408"},{"key":"e_1_3_1_152_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1514"},{"key":"e_1_3_1_153_2","article-title":"Pixel-BERT: Aligning image pixels with text by deep multi-modal transformers","author":"Huang Zhicheng","year":"2020","unstructured":"Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. 2020. Pixel-BERT: Aligning image pixels with text by deep multi-modal transformers. arXiv preprint arXiv:2004.00849 (2020).","journal-title":"arXiv preprint arXiv:2004.00849"},{"key":"e_1_3_1_154_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01003"},{"key":"e_1_3_1_155_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.7005"},{"key":"e_1_3_1_156_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_1_157_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.667"},{"key":"e_1_3_1_158_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.131"},{"key":"e_1_3_1_159_2","volume-title":"European Conference on Machine Learning (ECML) Workshops","author":"Chappuis Christel","year":"2021","unstructured":"Christel Chappuis, Sylvain Lobry, Benjamin Alexander Kellenberger, Bertrand Le Saux, and Devis Tuia. 2021. How to find a good image-text embedding for remote sensing visual question answering?. In European Conference on Machine Learning (ECML) Workshops."},{"key":"e_1_3_1_160_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00181"},{"key":"e_1_3_1_161_2","article-title":"Analyzing compositionality in visual question answering.","volume":"7","author":"Subramanian Sanjay","year":"2019","unstructured":"Sanjay Subramanian, Sameer Singh, and Matt Gardner. 2019. Analyzing compositionality in visual question answering. Advances in Neural Information Processing Systems 7 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_162_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00331"},{"key":"e_1_3_1_163_2","unstructured":"Paul Pu Liang Yiwei Lyu Xiang Fan Zetian Wu Yun Cheng Jason Wu Leslie Yufan Chen Peter Wu Michelle A Lee Yuke Zhu Ruslan Salakhutdinov and Louis-Philippe Morency. MultiBench: Multiscale Benchmarks for Multimodal Representation Learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1) ."},{"key":"e_1_3_1_164_2","volume-title":"Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)","author":"Shi Xingjian","unstructured":"Xingjian Shi, Jonas Mueller, Nick Erickson, Mu Li, and Alex Smola. Benchmarking multimodal AutoML for tabular data with text fields. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)."},{"key":"e_1_3_1_165_2","article-title":"Attentive explanations: Justifying decisions and pointing to the evidence","author":"Park Dong Huk","year":"2016","unstructured":"Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach. 2016. Attentive explanations: Justifying decisions and pointing to the evidence. arXiv preprint arXiv:1612.04757 (2016).","journal-title":"arXiv preprint arXiv:1612.04757"},{"key":"e_1_3_1_166_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00522"},{"key":"e_1_3_1_167_2","first-page":"1060","volume-title":"International Conference on Machine Learning","author":"Reed Scott","year":"2016","unstructured":"Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016. Generative adversarial text to image synthesis. In International Conference on Machine Learning. PMLR, 1060\u20131069."},{"key":"e_1_3_1_168_2","article-title":"The Caltech-UCSD Birds-200-2011 dataset","author":"Wah Catherine","year":"2011","unstructured":"Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. 2011. The Caltech-UCSD Birds-200-2011 dataset. California Institute of Technology (2011).","journal-title":"California Institute of Technology"},{"key":"e_1_3_1_169_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00143"},{"key":"e_1_3_1_170_2","first-page":"2768","volume-title":"INTERSPEECH","author":"Qu Leyuan","year":"2019","unstructured":"Leyuan Qu, Cornelius Weber, and Stefan Wermter. 2019. LipSound: Neural mel-spectrogram reconstruction for lip reading. In INTERSPEECH. 2768\u20132772."},{"key":"e_1_3_1_171_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2889052"},{"key":"e_1_3_1_172_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2015.2407694"},{"key":"e_1_3_1_173_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Ping Wei","year":"2018","unstructured":"Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller. 2018. Deep voice 3: Scaling text-to-speech with convolutional sequence learning. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_174_2","doi-asserted-by":"crossref","unstructured":"Jonathan Shen Ruoming Pang Ron J. Weiss Mike Schuster Navdeep Jaitly Zongheng Yang Zhifeng Chen Yu Zhang Yuxuan Wang Rj Skerrv-Ryan Rif A. Saurous Yannis Agiomyrgiannakis and Yonghui Wu. 2018. Natural tts synthesis by conditioning wavenet on mel spectrogram predictions. In IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP\u201918) . IEEE 4779\u20134783.","DOI":"10.1109\/ICASSP.2018.8461368"},{"key":"e_1_3_1_175_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7953127"},{"key":"e_1_3_1_176_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461856"},{"key":"e_1_3_1_177_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-1445"},{"key":"e_1_3_1_178_2","volume-title":"International Conference on Learning Representations","author":"Su Weijie","year":"2020","unstructured":"Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020. VL-BERT: Pre-training of generic visual-linguistic representations. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=SygXPaEYvH"},{"key":"e_1_3_1_179_2","article-title":"Representation learning for electronic health records","author":"Weng Wei-Hung","year":"2019","unstructured":"Wei-Hung Weng and Peter Szolovits. 2019. Representation learning for electronic health records. arXiv preprint arXiv:1909.09248 (2019).","journal-title":"arXiv preprint arXiv:1909.09248"},{"key":"e_1_3_1_180_2","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611976700.66"},{"key":"e_1_3_1_181_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2022.3197127"},{"issue":"2","key":"e_1_3_1_182_2","first-page":"699","article-title":"Predicting the survival of cancer patients with multimodal graph neural network","volume":"19","author":"Gao Jianliang","year":"2021","unstructured":"Jianliang Gao, Tengfei Lyu, Fan Xiong, Jianxin Wang, Weimao Ke, and Zhao Li. 2021. Predicting the survival of cancer patients with multimodal graph neural network. IEEE\/ACM Transactions on Computational Biology and Bioinformatics 19, 2 (2021), 699\u2013709.","journal-title":"IEEE\/ACM Transactions on Computational Biology and Bioinformatics"},{"key":"e_1_3_1_183_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1422"},{"key":"e_1_3_1_184_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Bahdanau Dzmitry","year":"2015","unstructured":"Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In Proceedings of the International Conference on Learning Representations, Yoshua Bengio and Yann LeCun (Eds.)."},{"key":"e_1_3_1_185_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2020.11.024"},{"key":"e_1_3_1_186_2","doi-asserted-by":"publisher","DOI":"10.5555\/2566972.2566993"},{"key":"e_1_3_1_187_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1208"},{"key":"e_1_3_1_188_2","doi-asserted-by":"publisher","DOI":"10.1038\/sdata.2016.35"},{"key":"e_1_3_1_189_2","unstructured":"Xintong Han. 2017. Fashion 200K Benchmark. https:\/\/github.com\/xthan\/fashion-200k. (2017). [Online; accessed 2017]."},{"key":"e_1_3_1_190_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2011.6130298"},{"key":"e_1_3_1_191_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33715-4_54"},{"key":"e_1_3_1_192_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01275"},{"key":"e_1_3_1_193_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.571"},{"key":"e_1_3_1_194_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123427"},{"key":"e_1_3_1_195_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.149"},{"key":"e_1_3_1_196_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00647"},{"key":"e_1_3_1_197_2","volume-title":"Visually Grounded Interaction and Language (ViGIL), NeurIPS 2019 Workshop","author":"Cangea Catalina","year":"2019","unstructured":"Catalina Cangea, Eugene Belilovsky, Pietro Li\u00f2, and Aaron C. Courville. 2019. VideoNavQA: Bridging the gap between visual and embodied question answering. In Visually Grounded Interaction and Language (ViGIL), NeurIPS 2019 Workshop."},{"key":"e_1_3_1_198_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.217"},{"key":"e_1_3_1_199_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"e_1_3_1_200_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICVGIP.2008.47"},{"key":"e_1_3_1_201_2","article-title":"MOSI: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos","author":"Zadeh Amir","year":"2016","unstructured":"Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency. 2016. MOSI: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos. arXiv preprint arXiv:1606.06259 (2016).","journal-title":"arXiv preprint arXiv:1606.06259"},{"key":"e_1_3_1_202_2","doi-asserted-by":"publisher","DOI":"10.1109\/FG47880.2020.00134"},{"key":"e_1_3_1_203_2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913509035"},{"key":"e_1_3_1_204_2","first-page":"1298","volume-title":"International Conference on Machine Learning","author":"Baevski Alexei","year":"2022","unstructured":"Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. 2022. Data2vec: A general framework for self-supervised learning in speech, vision and language. In International Conference on Machine Learning. PMLR, 1298\u20131312."},{"key":"e_1_3_1_205_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01519"},{"key":"e_1_3_1_206_2","first-page":"1","article-title":"Multimodal learning with graphs","author":"Ektefaie Yasha","year":"2023","unstructured":"Yasha Ektefaie, George Dasoulas, Ayush Noori, Maha Farhat, and Marinka Zitnik. 2023. Multimodal learning with graphs. Nature Machine Intelligence (2023), 1\u201311.","journal-title":"Nature Machine Intelligence"},{"key":"e_1_3_1_207_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00397"},{"key":"e_1_3_1_208_2","article-title":"MURAL: Multimodal, multitask retrieval across languages","author":"Jain Aashi","year":"2021","unstructured":"Aashi Jain, Mandy Guo, Krishna Srinivasan, Ting Chen, Sneha Kudugunta, Chao Jia, Yinfei Yang, and Jason Baldridge. 2021. MURAL: Multimodal, multitask retrieval across languages. arXiv preprint arXiv:2109.05125 (2021).","journal-title":"arXiv preprint arXiv:2109.05125"},{"key":"e_1_3_1_209_2","article-title":"Cross-view language modeling: Towards unified cross-lingual cross-modal pre-training","author":"Zeng Yan","year":"2022","unstructured":"Yan Zeng, Wangchunshu Zhou, Ao Luo, and Xinsong Zhang. 2022. Cross-view language modeling: Towards unified cross-lingual cross-modal pre-training. arXiv preprint arXiv:2206.00621 (2022).","journal-title":"arXiv preprint arXiv:2206.00621"},{"key":"e_1_3_1_210_2","first-page":"9694","article-title":"Align before fuse: Vision and language representation learning with momentum distillation","volume":"34","author":"Li Junnan","year":"2021","unstructured":"Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. Advances in Neural Information Processing Systems 34 (2021), 9694\u20139705.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_211_2","unstructured":"Xi Chen Xiao Wang Soravit Changpinyo AJ Piergiovanni Piotr Padlewski Daniel Salz Sebastian Goodman Adam Grycner Basil Mustafa Lucas Beyer Alexander Kolesnikov Joan Puigcerver Nan Ding Keran Rong Hassan Akbari Gaurav Mishra Linting Xue Ashish V. Thapliyal James Bradbury Weicheng Kuo Mojtaba Seyedhosseini Chao Jia Burcu Karagol Ayan Carlos Riquelme Andreas Steiner Anelia Angelova Xiaohua Zhai Neil Houlsby and Radu Soricut. 2022. Pali: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794 (2022)."},{"key":"e_1_3_1_212_2","first-page":"200","article-title":"Multimodal few-shot learning with frozen language models","volume":"34","author":"Tsimpoukelli Maria","year":"2021","unstructured":"Maria Tsimpoukelli, Jacob L. Menick, Serkan Cabi, S. M. Eslami, Oriol Vinyals, and Felix Hill. 2021. Multimodal few-shot learning with frozen language models. Advances in Neural Information Processing Systems 34 (2021), 200\u2013212.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_213_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01246"},{"key":"e_1_3_1_214_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11633-022-1369-5"},{"key":"e_1_3_1_215_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00146"},{"key":"e_1_3_1_216_2","first-page":"4904","volume-title":"International Conference on Machine Learning","author":"Jia Chao","year":"2021","unstructured":"Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning. PMLR, 4904\u20134916."},{"key":"e_1_3_1_217_2","doi-asserted-by":"crossref","unstructured":"Nanyi Fei Zhiwu Lu Yizhao Gao Guoxing Yang Yuqi Huo Jingyuan Wen Haoyu Lu Ruihua Song Xin Gao Tao Xiang Haoran Sun and Jiling Wen. 2022. Towards artificial general intelligence via a multimodal foundation model. Nature Communications 13 1 (2022) 3094.","DOI":"10.1038\/s41467-022-30761-2"},{"key":"e_1_3_1_218_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2021.3066303"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3617833","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3617833","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:57Z","timestamp":1750178277000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3617833"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,23]]},"references-count":217,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,3,31]]}},"alternative-id":["10.1145\/3617833"],"URL":"https:\/\/doi.org\/10.1145\/3617833","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,23]]},"assertion":[{"value":"2023-01-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-08-10","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-10-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}