{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T14:48:53Z","timestamp":1773326933008,"version":"3.50.1"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2024,5,10]],"date-time":"2024-05-10T00:00:00Z","timestamp":1715299200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2024,5,31]]},"abstract":"<jats:p>Lying in the cross-section of computer vision and natural language processing, vision language models are capable of processing images and text at once. These models are helpful in various tasks: text generation from image and vice versa, image-text retrieval, or visual navigation. Besides building a model trained on a dataset for a task, people also study general-purpose models to utilize many datasets for multitasks. Their two primary applications are image captioning and visual question answering. For English, large datasets and foundation models are already abundant. However, for Vietnamese, they are still limited. To expand the language range, this work proposes a pretrained general-purpose image-text model named VisualRoBERTa. A dataset of 600k images with captions (translated MS COCO 2017 from English to Vietnamese) is introduced to pretrain VisualRoBERTa. The model\u2019s architecture is built using Convolutional Neural Network and Transformer blocks. Fine-tuning VisualRoBERTa shows promising results on the ViVQA dataset with 34.49% accuracy, 0.4173 BLEU 4, and 0.4390 RougeL (in visual question answering task), and best outcomes on the sViIC dataset with 0.6685 BLEU 4, 0.6320 RougeL (in image captioning task).<\/jats:p>","DOI":"10.1145\/3654796","type":"journal-article","created":{"date-parts":[[2024,3,30]],"date-time":"2024-03-30T09:24:01Z","timestamp":1711790641000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["A Novel Pretrained General-purpose Vision Language Model for the Vietnamese Language"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8071-7962","authenticated-orcid":false,"given":"Dinh Anh","family":"Vu","sequence":"first","affiliation":[{"name":"University of Science and Technology of Hanoi, Hanoi, Viet Nam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0570-4619","authenticated-orcid":false,"given":"Quang Nhat Minh","family":"Pham","sequence":"additional","affiliation":[{"name":"Aimesoft JSC, Hanoi Viet Nam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7140-6882","authenticated-orcid":false,"given":"Giang Son","family":"Tran","sequence":"additional","affiliation":[{"name":"University of Science and Technology of Hanoi, Hanoi, Viet Nam"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,5,10]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00904"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_2_4_2","article-title":"On the opportunities and risks of foundation models","author":"Bommasani Rishi","year":"2021","unstructured":"Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jared Quincy Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark Krass, Ranjay Krishna, Rohith Kuditipudi, Ananya Kumar, Faisal Ladhak, Mina Lee, Tony Lee, Jure Leskovec, Isabelle Levent, Xiang Lisa Li, Xuechen Li, Tengyu Ma, Ali Malik, Christopher D. Manning, Suvir Mirchandani, Eric Mitchell, Zanele Munyikwa, Suraj Nair, Avanika Narayan, Deepak Narayanan, Ben Newman, Allen Nie, Juan Carlos Niebles, Hamed Nilforoshan, Julian Nyarko, Giray Ogut, Laurel Orr, Isabel Papadimitriou, Joon Sung Park, Chris Piech, Eva Portelance, Christopher Potts, Aditi Raghunathan, Rob Reich, Hongyu Ren, Frieda Rong, Yusuf Roohani, Camilo Ruiz, Jack Ryan, Christopher R\u00e9, Dorsa Sadigh, Shiori Sagawa, Keshav Santhanam, Andy Shih, Krishnan Srinivasan, Alex Tamkin, Rohan Taori, Armin W. Thomas, Florian Tram\u00e8r, Rose E. Wang, William Wang, Bohan Wu, Jiajun Wu, Yuhuai Wu, Sang Michael Xie, Michihiro Yasunaga, Jiaxuan You, Matei Zaharia, Michael Zhang, Tianyi Zhang, Xikun Zhang, Yuhui Zhang, Lucia Zheng, Kaitlyn Zhou, and Percy Liang. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv: Arxiv-2108.07258 (2021).","journal-title":"arXiv preprint arXiv: Arxiv-2108.07258"},{"key":"e_1_3_2_5_2","article-title":"VLP: A survey on vision-language pre-training","author":"Chen Feilong","year":"2022","unstructured":"Feilong Chen, Duzhen Zhang, Minglun Han, Xiuyi Chen, Jing Shi, Shuang Xu, and Bo Xu. 2022. VLP: A survey on vision-language pre-training. arXiv preprint arXiv: Arxiv-2202.09061 (2022).","journal-title":"arXiv preprint arXiv: Arxiv-2202.09061"},{"key":"e_1_3_2_6_2","article-title":"PaLI: A jointly-scaled multilingual language-image model","author":"Chen Xi","year":"2022","unstructured":"Xi Chen, Xiao Wang, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut. 2022. PaLI: A jointly-scaled multilingual language-image model. arXiv preprint arXiv: Arxiv-2209.06794 (2022).","journal-title":"arXiv preprint arXiv: Arxiv-2209.06794"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_7"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.369"},{"key":"e_1_3_2_9_2","article-title":"A survey of vision-language pre-trained models","author":"Du Yifan","year":"2022","unstructured":"Yifan Du, Zikang Liu, Junyi Li, and Wayne Xin Zhao. 2022. A survey of vision-language pre-trained models. arXiv preprint arXiv:2202.10936 (2022).","journal-title":"arXiv preprint arXiv:2202.10936"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00146"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.670"},{"key":"e_1_3_2_12_2","article-title":"Pixel-BERT: Aligning image pixels with text by deep multi-modal transformers","author":"Huang Zhicheng","year":"2020","unstructured":"Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. 2020. Pixel-BERT: Aligning image pixels with text by deep multi-modal transformers. arXiv preprint arXiv:2004.00849 (2020).","journal-title":"arXiv preprint arXiv:2004.00849"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00686"},{"key":"e_1_3_2_14_2","series-title":"Proceedings of Machine Learning Research","first-page":"5583","volume-title":"Proceedings of the 38th International Conference on Machine Learning (ICML\u20192","volume":"139","author":"Kim Wonjae","year":"2021","unstructured":"Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. ViLT: Vision-and-language transformer without convolution or region supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML\u20192(Proceedings of Machine Learning Research, Vol. 139), Marina Meila and Tong Zhang (Eds.). PMLR, 5583\u20135594. Retrieved from http:\/\/proceedings.mlr.press\/v139\/kim21k.html"},{"key":"e_1_3_2_15_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations (ICLR\u201915)","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the 3rd International Conference on Learning Representations (ICLR\u201915), Yoshua Bengio and Yann LeCun (Eds.). Retrieved from http:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_2_16_2","unstructured":"Ranjay Krishna Yuke Zhu Oliver Groth Justin Johnson Kenji Hata Joshua Kravitz Stephanie Chen Yannis Kalantidis Li-Jia Li David A. Shamma Michael Bernstein and Li Fei-Fei. 2016. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Retrieved from https:\/\/arxiv.org\/abs\/1602.07332"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","unstructured":"Quan Hoang Lam Quang Duy Le Kiet Van Nguyen and Ngan Luu-Thuy Nguyen. 2020. UIT-ViIC: A Dataset for the First Evaluation on Vietnamese Image Captioning. DOI:10.48550\/ARXIV.2002.00175","DOI":"10.48550\/ARXIV.2002.00175"},{"key":"e_1_3_2_18_2","volume-title":"Proceedings of the 8th International Workshop on Vietnamese Language and Speech Processing","author":"Le Thao Minh","year":"2021","unstructured":"Thao Minh Le, Long Hoang Dang, Thanh-Son Nguyen, Thi Minh Huyen Nguyen, and Xuan-Son Vu. 2021. VLSP 2021 - VieCap4H challenge: Automatic image caption generation for healthcare domain in Vietnamese. In Proceedings of the 8th International Workshop on Vietnamese Language and Speech Processing."},{"key":"e_1_3_2_19_2","unstructured":"Thanh V. Le. 2021. Pretrained GPT-2 on Vietnamese news. Retrieved from https:\/\/huggingface.co\/imthanhlv\/gpt2news"},{"key":"e_1_3_2_20_2","first-page":"12888","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning. PMLR, 12888\u201312900."},{"key":"e_1_3_2_21_2","article-title":"VisualBERT: A simple and performant baseline for vision and language","author":"Li Liunian Harold","year":"2019","unstructured":"Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. VisualBERT: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557 (2019).","journal-title":"arXiv preprint arXiv:1908.03557"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"e_1_3_2_23_2","first-page":"74","volume-title":"Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out. Association for Computational Linguistics, 74\u201381. Retrieved from https:\/\/aclanthology.org\/W04-1013"},{"key":"e_1_3_2_24_2","article-title":"InterBERT: Vision-and-language interaction for multi-modal pretraining","author":"Lin Junyang","year":"2020","unstructured":"Junyang Lin, An Yang, Yichang Zhang, Jie Liu, Jingren Zhou, and Hongxia Yang. 2020. InterBERT: Vision-and-language interaction for multi-modal pretraining. arXiv preprint arXiv:2003.13198 (2020).","journal-title":"arXiv preprint arXiv:2003.13198"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_26_2","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv: Arxiv-1907.11692 (2019).","journal-title":"arXiv preprint arXiv: Arxiv-1907.11692"},{"key":"e_1_3_2_27_2","article-title":"Vision-and-language pretrained models: A survey","author":"Long Siqu","year":"2022","unstructured":"Siqu Long, Feiqi Cao, Soyeon Caren Han, and Haiqing Yang. 2022. Vision-and-language pretrained models: A survey. arXiv preprint arXiv:2204.07356 (2022).","journal-title":"arXiv preprint arXiv:2204.07356"},{"key":"e_1_3_2_28_2","article-title":"VilBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks","volume":"32","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. VilBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Adv. Neural Inf. Process. Syst. 32 (2019).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1.11688"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.findings-emnlp.92"},{"key":"e_1_3_2_31_2","article-title":"Vision-and-language pretraining","author":"Nguyen Thong","year":"2022","unstructured":"Thong Nguyen, Cong-Duy Nguyen, Xiaobao Wu, and Anh Tuan Luu. 2022. Vision-and-language pretraining. arXiv preprint arXiv: Arxiv-2207.01772 (2022).","journal-title":"arXiv preprint arXiv: Arxiv-2207.01772"},{"key":"e_1_3_2_32_2","volume-title":"Proceedings of the 23rd Annual Conference of the International Speech Communication Association: Show and Tell (INTERSPEECH\u201922)","author":"Nguyen Thien Hai","year":"2022","unstructured":"Thien Hai Nguyen, Tuan-Duy H. Nguyen, Duy Phung, Duy Tran-Cong Nguyen, Hieu Minh Tran, Manh Luong, Tin Duy Vo, Hung Hai Bui, Dinh Phung, and Dat Quoc Nguyen. 2022. A Vietnamese-English neural machine translation system. In Proceedings of the 23rd Annual Conference of the International Speech Communication Association: Show and Tell (INTERSPEECH\u201922)."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00397"},{"key":"e_1_3_2_34_2","volume-title":"Proceedings of the Neural Information Processing Systems Conference (NIPS\u201911)","author":"Ordonez Vicente","year":"2011","unstructured":"Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. 2011. Im2Text: Describing images using 1 Million captioned photographs. In Proceedings of the Neural Information Processing Systems Conference (NIPS\u201911)."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_2_36_2","article-title":"How do vision transformers work? In","author":"Park Namuk","year":"2022","unstructured":"Namuk Park and Songkuk Kim. 2022. How do vision transformers work? In Proceedings of the International Conference on Learning Representations (ICLR\u201922).","journal-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201922)"},{"key":"e_1_3_2_37_2","first-page":"8024","volume-title":"Advances in Neural Information Processing Systems 32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch\u00e9-Buc, E. Fox, and R. Garnett (Eds.). Curran Associates, Inc., 8024\u20138035. Retrieved from http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-srw.18"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.2478\/pralin-2018-0002"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","unstructured":"Alec Radford Jong Wook Kim Chris Hallacy Aditya Ramesh Gabriel Goh Sandhini Agarwal Girish Sastry Amanda Askell Pamela Mishkin Jack Clark Gretchen Krueger and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. DOI:10.48550\/ARXIV.2103.00020","DOI":"10.48550\/ARXIV.2103.00020"},{"key":"e_1_3_2_41_2","first-page":"140:1\u2013140:67","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21 (2020), 140:1\u2013140:67. Retrieved from http:\/\/jmlr.org\/papers\/v21\/20-074.html","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_2_42_2","article-title":"Exploring models and data for image question answering","volume":"28","author":"Ren Mengye","year":"2015","unstructured":"Mengye Ren, Ryan Kiros, and Richard Zemel. 2015. Exploring models and data for image question answering. Adv. Neural Inf. Process. Syst. 28 (2015). Retrieved from https:\/\/www.cs.toronto.edu\/mren\/research\/imageqa\/data\/cocoqa\/","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_2_43_2","article-title":"LAION-400M: Open dataset of CLIP-Filtered 400 Million image-text pairs","author":"Schuhmann Christoph","year":"2021","unstructured":"Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. 2021. LAION-400M: Open dataset of CLIP-Filtered 400 Million image-text pairs. arXiv preprint arXiv: Arxiv-2111.02114 (2021).","journal-title":"arXiv preprint arXiv: Arxiv-2111.02114"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1238"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1238"},{"key":"e_1_3_2_46_2","article-title":"Vl-BERT: Pre-training of generic visual-linguistic representations","author":"Su Weijie","year":"2019","unstructured":"Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2019. Vl-BERT: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530 (2019).","journal-title":"arXiv preprint arXiv:1908.08530"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-2034"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1644"},{"key":"e_1_3_2_49_2","article-title":"Recent advances and trends in multimodal deep learning: A review","author":"Summaira Jabeen","year":"2021","unstructured":"Jabeen Summaira, Xi Li, Amin Muhammad Shoib, Songyuan Li, and Jabbar Abdul. 2021. Recent advances and trends in multimodal deep learning: A review. arXiv preprint arXiv:2105.11087 (2021).","journal-title":"arXiv preprint arXiv:2105.11087"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1514"},{"key":"e_1_3_2_51_2","first-page":"683","volume-title":"Proceedings of the 35th Pacific Asia Conference on Language, Information and Computation","author":"Tran Khanh Quoc","year":"2021","unstructured":"Khanh Quoc Tran, An Trong Nguyen, An Tran-Hoai Le, and Kiet Van Nguyen. 2021. ViVQA: Vietnamese visual question answering. In Proceedings of the 35th Pacific Asia Conference on Language, Information and Computation. Association for Computational Lingustics, 683\u2013691. Retrieved from https:\/\/aclanthology.org\/2021.paclic-1.72"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","unstructured":"Nguyen Luong Tran Duong Minh Le and Dat Quoc Nguyen. 2021. BARTpho: Pre-trained Sequence-to-Sequence Models for Vietnamese. DOI:10.48550\/ARXIV.2109.09701","DOI":"10.48550\/ARXIV.2109.09701"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_2_54_2","article-title":"MiniVLM: A smaller and faster vision-language model","author":"Wang Jianfeng","year":"2020","unstructured":"Jianfeng Wang, Xiaowei Hu, Pengchuan Zhang, Xiujun Li, Lijuan Wang, Lei Zhang, Jianfeng Gao, and Zicheng Liu. 2020. MiniVLM: A smaller and faster vision-language model. arXiv preprint arXiv: Arxiv-2012.06946 (2020).","journal-title":"arXiv preprint arXiv: Arxiv-2012.06946"},{"key":"e_1_3_2_55_2","article-title":"SimVLM: Simple visual language model pretraining with weak supervision","author":"Wang Zirui","year":"2022","unstructured":"Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. 2022. SimVLM: Simple visual language model pretraining with weak supervision. In Proceedings of the International Conference on Learning Representations (ICLR\u201922).","journal-title":"Proceedings of the International Conference on Learning Representations (ICLR\u201922)"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","unstructured":"Thomas Wolf Lysandre Debut Victor Sanh Julien Chaumond Clement Delangue Anthony Moi Perric Cistac Clara Ma Yacine Jernite Julien Plu Canwen Xu Teven Le Scao Sylvain Gugger Mariama Drame Quentin Lhoest and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for Computational Linguistics 38\u201345. 10.18653\/v1\/2020.emnlp-demos.6","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","unstructured":"Zhibiao Wu and Martha Palmer. 1994. Verb semantics and lexical selection. DOI:10.48550\/ARXIV.CMP-LG\/9406033","DOI":"10.48550\/ARXIV.CMP-LG\/9406033"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","unstructured":"Rowan Zellers Yonatan Bisk Ali Farhadi and Yejin Choi. 2019. From recognition to cognition: Visual commonsense reasoning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 6720\u20136731. 10.1109\/CVPR.2019.00688","DOI":"10.1109\/CVPR.2019.00688"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.7005"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00414"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3654796","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3654796","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:06:09Z","timestamp":1750291569000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3654796"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,10]]},"references-count":59,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,5,31]]}},"alternative-id":["10.1145\/3654796"],"URL":"https:\/\/doi.org\/10.1145\/3654796","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,10]]},"assertion":[{"value":"2023-05-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-03-24","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-05-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}