{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T17:23:45Z","timestamp":1777051425711,"version":"3.51.4"},"reference-count":158,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T00:00:00Z","timestamp":1759190400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T00:00:00Z","timestamp":1759190400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Curtin University"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Multimed Info Retr"],"published-print":{"date-parts":[[2025,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>The primary goal of multimodal visual representation learning is to generate implicit information that effectively represents multimodal information by exploring the commonalities and characteristics between different modalities. This research report will discuss currently widely used advanced methods in the field of multimodal visual representation learning. This article will discuss these methods in the following order, culminating in multimodal visual learning: (1) pre-trained visual representation learning, (2) generative visual representation learning, (3) contrastive multimodal visual representation learning, and (4) image-text multimodal visual representation learning methods. Each element provides useful clues that ultimately lead to multimodal visual learning. Pre-trained visual representation learning refers to the application of supervised pre-training models in visual representation learning, while generative visual representation learning uses generative models to learn feature representations that can integrate multimodal information. Contrastive multimodal visual representation learning uses contrastive learning methods to compare similar and dissimilar sample pairs, learning feature representations in a self-supervised manner. Image-text multimodal visual representation learning methods, on the other hand, attempt to enhance the capabilities of visual representation learning by fusing visual information (such as images) with textual information. This review report will explain the above research background, the classification of different research methods, commonly used evaluation methods , and future development trends.<\/jats:p>","DOI":"10.1007\/s13735-025-00382-8","type":"journal-article","created":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T17:46:42Z","timestamp":1759254402000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["A Comprehensive Review of Multimodal Visual Representation Learning: Tracing the Evolution from CNNs to Transformers and Beyond"],"prefix":"10.1007","volume":"14","author":[{"given":"Dong","family":"Zhang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"W. K.","family":"Wong","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"I. M.","family":"Chew","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,9,30]]},"reference":[{"issue":"11","key":"382_CR1","doi-asserted-by":"publisher","first-page":"65","DOI":"10.1109\/35.41402","volume":"27","author":"BP Yuhas","year":"1989","unstructured":"Yuhas BP, Goldstein MH, Sejnowski TJ (1989) Integration of acoustic and visual speech signals using neural networks. IEEE Commun Mag 27(11):65\u201371","journal-title":"IEEE Commun Mag"},{"issue":"3","key":"382_CR2","doi-asserted-by":"publisher","first-page":"251","DOI":"10.1080\/00401706.1991.10484833","volume":"33","author":"BH Juang","year":"1991","unstructured":"Juang BH, Rabiner LR (1991) Hidden markov models for speech recognition. Technometrics 33(3):251\u2013272","journal-title":"Technometrics"},{"issue":"2","key":"382_CR3","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1109\/TPAMI.2018.2798607","volume":"41","author":"T Baltru\u0161aitis","year":"2018","unstructured":"Baltru\u0161aitis T, Ahuja C, Morency L-P (2018) Multimodal machine learning: a survey and taxonomy. IEEE Trans Pattern Anal Mach Intell 41(2):423\u2013443","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"8","key":"382_CR4","doi-asserted-by":"publisher","first-page":"1798","DOI":"10.1109\/TPAMI.2013.50","volume":"35","author":"Y Bengio","year":"2013","unstructured":"Bengio Y, Courville A, Vincent P (2013) Representation learning: a review and new perspectives. IEEE Trans Pattern Anal Mach Intell 35(8):1798\u20131828","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"382_CR5","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","volume":"60","author":"DG Lowe","year":"2004","unstructured":"Lowe DG (2004) Distinctive image features from scale-invariant keypoints. Int J Comput Vision 60:91\u2013110","journal-title":"Int J Comput Vision"},{"key":"382_CR6","doi-asserted-by":"crossref","unstructured":"Dalal N, Triggs B (2005) Histograms of oriented gradients for human detection. In: 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905), vol. 1, pp. 886\u2013893","DOI":"10.1109\/CVPR.2005.177"},{"issue":"7","key":"382_CR7","doi-asserted-by":"publisher","first-page":"1527","DOI":"10.1162\/neco.2006.18.7.1527","volume":"18","author":"GE Hinton","year":"2006","unstructured":"Hinton GE, Osindero S, Teh Y-W (2006) A fast learning algorithm for deep belief nets. Neural Comput 18(7):1527\u20131554","journal-title":"Neural Comput"},{"issue":"1","key":"382_CR8","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1109\/TBDATA.2018.2850013","volume":"6","author":"D Zhang","year":"2018","unstructured":"Zhang D, Yin J, Zhu X, Zhang C (2018) Network representation learning: a survey. IEEE transactions on Big Data 6(1):3\u201328","journal-title":"IEEE transactions on Big Data"},{"issue":"4","key":"382_CR9","doi-asserted-by":"publisher","first-page":"415","DOI":"10.1016\/j.clinbiomech.2004.01.005","volume":"19","author":"A Daffertshofer","year":"2004","unstructured":"Daffertshofer A, Lamoth CJ, Meijer OG, Beek PJ (2004) Pca in studying coordination and variability: a tutorial. Clin Biomech 19(4):415\u2013428","journal-title":"Clin Biomech"},{"issue":"4","key":"382_CR10","doi-asserted-by":"publisher","first-page":"193","DOI":"10.1007\/BF00344251","volume":"36","author":"K Fukushima","year":"1980","unstructured":"Fukushima K (1980) Neocognitron: a self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biol Cybern 36(4):193\u2013202","journal-title":"Biol Cybern"},{"issue":"11","key":"382_CR11","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","volume":"86","author":"Y LeCun","year":"1998","unstructured":"LeCun Y, Bottou L, Bengio Y, Haffner P (1998) Gradient-based learning applied to document recognition. Proc IEEE 86(11):2278\u20132324","journal-title":"Proc IEEE"},{"issue":"6","key":"382_CR12","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1145\/3065386","volume":"60","author":"A Krizhevsky","year":"2017","unstructured":"Krizhevsky A, Sutskever I, Hinton GE (2017) Imagenet classification with deep convolutional neural networks. Commun ACM 60(6):84\u201390","journal-title":"Commun ACM"},{"key":"382_CR13","unstructured":"Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556"},{"key":"382_CR14","unstructured":"Schmid C, Soatto S, Tomasi C (2005) Conference on Computer Vision and Pattern Recognition. IEEE Computer Society"},{"key":"382_CR15","doi-asserted-by":"crossref","unstructured":"Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z (2016) Rethinking the inception architecture for computer vision. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2818\u20132826","DOI":"10.1109\/CVPR.2016.308"},{"key":"382_CR16","doi-asserted-by":"crossref","unstructured":"Szegedy C, Ioffe S, Vanhoucke V, Alemi A (2017) Inception-v4, inception-resnet and the impact of residual connections on learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 31","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"382_CR17","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"382_CR18","doi-asserted-by":"crossref","unstructured":"Huang G, Liu Z, Van Der\u00a0Maaten L, Weinberger KQ (2017) Densely connected convolutional networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4700\u20134708","DOI":"10.1109\/CVPR.2017.243"},{"key":"382_CR19","unstructured":"Iandola FN, Han S, Moskewicz MW, Ashraf K, Dally WJ, Keutzer K (2016) Squeezenet: Alexnet-level accuracy with 50x fewer parameters and$$<$$ 0.5 mb model size. arXiv preprint arXiv:1602.07360"},{"key":"382_CR20","unstructured":"Howard AG, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Andreetto M, Adam H (2017) Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861"},{"key":"382_CR21","doi-asserted-by":"crossref","unstructured":"Zhang X, Zhou X, Lin M, Sun J (2018) Shufflenet: An extremely efficient convolutional neural network for mobile devices. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6848\u20136856","DOI":"10.1109\/CVPR.2018.00716"},{"key":"382_CR22","doi-asserted-by":"crossref","unstructured":"Hu J, Shen L, Sun G (2018) Squeeze-and-excitation networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7132\u20137141","DOI":"10.1109\/CVPR.2018.00745"},{"key":"382_CR23","unstructured":"Tan M, Le Q (2019) Efficientnet: Rethinking model scaling for convolutional neural networks. In: International Conference on Machine Learning, pp. 6105\u20136114 . PMLR"},{"key":"382_CR24","doi-asserted-by":"crossref","unstructured":"Radosavovic I, Kosaraju RP, Girshick R, He K, Doll\u00e1r P (2020) Designing network design spaces. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 10428\u201310436","DOI":"10.1109\/CVPR42600.2020.01044"},{"key":"382_CR25","unstructured":"Brock A, De S, Smith SL, Simonyan K (2021) High-performance large-scale image recognition without normalization. In: International Conference on Machine Learning, pp. 1059\u20131071 . PMLR"},{"key":"382_CR26","doi-asserted-by":"crossref","unstructured":"Liu Z, Mao H, Wu C-Y, Feichtenhofer C, Darrell T, Xie S (2022) A convnet for the 2020s. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 11976\u201311986","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"382_CR27","doi-asserted-by":"crossref","unstructured":"Woo S, Debnath S, Hu R, Chen X, Liu Z, Kweon IS, Xie S (2023) Convnext v2: Co-designing and scaling convnets with masked autoencoders. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 16133\u201316142","DOI":"10.1109\/CVPR52729.2023.01548"},{"key":"382_CR28","doi-asserted-by":"crossref","unstructured":"Qin D, Leichner C, Delakis M, Fornoni M, Luo S, Yang F, Wang W, Banbury C, Ye C, Akin B (2024) et al.: Mobilenetv4: Universal models for the mobile ecosystem. In: European Conference on Computer Vision, pp. 78\u201396 Springer","DOI":"10.1007\/978-3-031-73661-2_5"},{"key":"382_CR29","doi-asserted-by":"crossref","unstructured":"Zhai J, Zhang S, Chen J, He Q (2018) Autoencoder and its various variants. In: 2018 IEEE International Conference on Systems, Man, and Cybernetics (SMC), pp. 415\u2013419","DOI":"10.1109\/SMC.2018.00080"},{"key":"382_CR30","unstructured":"Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2014) Generative adversarial nets. Advances in neural information processing systems 27"},{"issue":"9","key":"382_CR31","doi-asserted-by":"publisher","first-page":"10850","DOI":"10.1109\/TPAMI.2023.3261988","volume":"45","author":"F-A Croitoru","year":"2023","unstructured":"Croitoru F-A, Hondru V, Ionescu RT, Shah M (2023) Diffusion models in vision: a survey. IEEE Trans Pattern Anal Mach Intell 45(9):10850\u201310869","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"382_CR32","unstructured":"Ng A (2011) et al.: Sparse autoencoder. CS294A Lecture notes 72(2011), 1\u201319"},{"key":"382_CR33","unstructured":"Lin X, Zhu C, Zhang Q, Liu Y (2016) 3d keypoint detection based on deep neural network with sparse autoencoder. arXiv preprint arXiv:1605.00129"},{"key":"382_CR34","doi-asserted-by":"crossref","unstructured":"Meng Q, Catchpoole D, Skillicom D, Kennedy PJ(2017) Relational autoencoder for feature extraction. In: 2017 International Joint Conference on Neural Networks (IJCNN), pp. 364\u2013371","DOI":"10.1109\/IJCNN.2017.7965877"},{"key":"382_CR35","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2020.103411","volume":"105","author":"N An","year":"2020","unstructured":"An N, Ding H, Yang J, Au R, Ang TF (2020) Deep ensemble learning for alzheimer\u2019s disease classification. J Biomed Inform 105:103411","journal-title":"J Biomed Inform"},{"key":"382_CR36","doi-asserted-by":"crossref","unstructured":"Vincent P, Larochelle H, Bengio Y, Manzagol P-A (2008) Extracting and composing robust features with denoising autoencoders. In: Proceedings of the 25th International Conference on Machine Learning, pp. 1096\u20131103","DOI":"10.1145\/1390156.1390294"},{"key":"382_CR37","doi-asserted-by":"crossref","unstructured":"Gidaris S, Komodakis N (2019) Generating classification weights with gnn denoising autoencoders for few-shot learning. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 21\u201330","DOI":"10.1109\/CVPR.2019.00011"},{"issue":"4","key":"382_CR38","doi-asserted-by":"publisher","first-page":"1017","DOI":"10.1109\/TCYB.2016.2536638","volume":"47","author":"B Du","year":"2016","unstructured":"Du B, Xiong W, Wu J, Zhang L, Zhang L, Tao D (2016) Stacked convolutional denoising auto-encoders for feature representation. IEEE transactions on cybernetics 47(4):1017\u20131027","journal-title":"IEEE transactions on cybernetics"},{"key":"382_CR39","doi-asserted-by":"crossref","unstructured":"He K, Chen X, Xie S, Li Y, Doll\u00e1r P, Girshick R (2022) Masked autoencoders are scalable vision learners. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 16000\u201316009","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"382_CR40","unstructured":"Chen J, Hu M, Li B, Elhoseiny M (2022) Efficient self-supervised vision pretraining with local masked reconstruction. arXiv preprint arXiv:2206.00790"},{"key":"382_CR41","unstructured":"Salah R, Vincent P, Muller X, Gloro X, Bengio Y (2011) Contractive auto-encoders: Explicit invariance during feature extraction. In: Proc. of the 28th International Conference on Machine Learning, pp. 833\u2013840"},{"key":"382_CR42","doi-asserted-by":"crossref","unstructured":"Ganguli S, Iyer CK, Pandey V (2022) Reachability embeddings: Scalable self-supervised representation learning from mobility trajectories for multimodal geospatial computer vision. In: 2022 23rd IEEE International Conference on Mobile Data Management (MDM), pp. 44\u201353 . IEEE","DOI":"10.1109\/MDM55031.2022.00028"},{"key":"382_CR43","unstructured":"Kingma DP, Welling M (2013) Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114"},{"key":"382_CR44","unstructured":"Sohn K, Lee H, Yan X (2015) Learning structured output representation using deep conditional generative models. Advances in neural information processing systems 28"},{"key":"382_CR45","unstructured":"Louizos C, Swersky K, Li Y, Welling M, Zemel R (2015) The variational fair autoencoder. arXiv preprint arXiv:1511.00830"},{"key":"382_CR46","unstructured":"Lopez R, Regier J, Jordan MI, Yosef N (2018) Information constraints on auto-encoding variational bayes. Advances in neural information processing systems 31"},{"key":"382_CR47","unstructured":"Ramachandra G (2017) Least square variational bayesian autoencoder with regularization. arXiv preprint arXiv:1707.03134"},{"key":"382_CR48","unstructured":"Chen X, Kingma DP, Salimans T, Duan Y, Dhariwal P, Schulman J, Sutskever I, Abbeel P (2016) Variational lossy autoencoder. arXiv preprint arXiv:1611.02731"},{"key":"382_CR49","doi-asserted-by":"publisher","DOI":"10.1016\/j.neuroimage.2023.119892","volume":"268","author":"G Mart\u00ed-Juan","year":"2023","unstructured":"Mart\u00ed-Juan G, Lorenzi M, Piella G, Initiative ADN et al (2023) Mc-rvae: multi-channel recurrent variational autoencoder for multimodal alzheimer\u2019s disease progression modelling. Neuroimage 268:119892","journal-title":"Neuroimage"},{"key":"382_CR50","doi-asserted-by":"crossref","unstructured":"Cai L, Gao H, Ji S (2019) Multi-stage variational auto-encoders for coarse-to-fine image generation. In: Proceedings of the 2019 SIAM International Conference on Data Mining, pp. 630\u2013638","DOI":"10.1137\/1.9781611975673.71"},{"key":"382_CR51","doi-asserted-by":"crossref","unstructured":"Liu AH, Jin S, Lai C-IJ, Rouditchenko A, Oliva A, Glass J (2021) Cross-modal discrete representation learning. arXiv preprint arXiv:2106.05438","DOI":"10.18653\/v1\/2022.acl-long.215"},{"key":"382_CR52","unstructured":"Razavi A, Oord A, Vinyals O (2019) Generating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems 32"},{"key":"382_CR53","doi-asserted-by":"crossref","unstructured":"Lygerakis F, Rueckert E (2023) Cr-vae: Contrastive regularization on variational autoencoders for preventing posterior collapse. In: 2023 7th Asian Conference on Artificial Intelligence Technology (ACAIT), pp. 427\u2013437","DOI":"10.1109\/ACAIT60137.2023.10528405"},{"key":"382_CR54","unstructured":"Goodfellow IJ, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2014) Generative adversarial networks, june 2014. arXiv preprint arXiv:1406.2661"},{"key":"382_CR55","unstructured":"Denton EL, Chintala S, Fergus R, et al (2015) Deep generative image models using a laplacian pyramid of adversarial networks. Advances in neural information processing systems 28"},{"key":"382_CR56","unstructured":"Radford A, Metz L, Chintala S (2015) Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434"},{"key":"382_CR57","unstructured":"Salimans T, Goodfellow I, Zaremba W, Cheung V, Radford A, Chen X (2016) Improved techniques for training gans. Advances in neural information processing systems 29"},{"key":"382_CR58","doi-asserted-by":"crossref","unstructured":"Lu S, Dong Z, Cai D, Fang F, Zhao D (2023) Mim-gan-based anomaly detection for multivariate time series data. In: 2023 IEEE 98th Vehicular Technology Conference (VTC2023-Fall), pp. 1\u20137","DOI":"10.1109\/VTC2023-Fall60731.2023.10333517"},{"key":"382_CR59","unstructured":"Wu J, Zhang C, Xue T, Freeman B, Tenenbaum J (2016) Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. Advances in neural information processing systems 29"},{"key":"382_CR60","doi-asserted-by":"crossref","unstructured":"Tang W, Li G, Bao X, Nian F, Li T(2020) Mscgan: Multi-scale conditional generative adversarial networks for person image generation. In: 2020 Chinese Control And Decision Conference (CCDC), pp. 1440\u20131445","DOI":"10.1109\/CCDC49329.2020.9164755"},{"key":"382_CR61","unstructured":"Sun H, Zhu T, Chang W, Zhou W (2023) Generative adversarial networks unlearning. arXiv preprint arXiv:2308.09881"},{"key":"382_CR62","doi-asserted-by":"crossref","unstructured":"Spandana C, Srisurya IV, Priyadharshini A, Krithika S, Nandhini SA, Kumar RP, Mohan GB (2023) Underwater image enhancement and restoration using cycle gan. In: International Conference On Innovative Computing And Communication, pp. 99\u2013110","DOI":"10.1007\/978-981-99-3010-4_9"},{"key":"382_CR63","unstructured":"Mirza M, Osindero S (2014) Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784"},{"key":"382_CR64","unstructured":"Reed SE, Akata Z, Mohan S, Tenka S, Schiele B, Lee H (2016) Learning what and where to draw. Advances in neural information processing systems 29"},{"key":"382_CR65","doi-asserted-by":"crossref","unstructured":"Zhang H, Xu T, Li H, Zhang S, Wang X, Huang X, Metaxas DN(2017) Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 5907\u20135915","DOI":"10.1109\/ICCV.2017.629"},{"key":"382_CR66","doi-asserted-by":"crossref","unstructured":"Bourou A, Boyer T, Gheisari M, Daupin K, Dubreuil V, De\u00a0Thonel A, Mezger V, Genovesio A (2024) Phendiff: Revealing subtle phenotypes with diffusion models in real images. In: International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 358\u2013367","DOI":"10.1007\/978-3-031-72384-1_34"},{"key":"382_CR67","doi-asserted-by":"crossref","unstructured":"Madhubalan A, Gautam A, Tiwary P (2024) Blender-gan: Multi-target conditional generative adversarial network for novel class synthetic data generation. In: 2024 International Conference on Smart Applications, Communications and Networking (SmartNets), pp. 1\u20137","DOI":"10.1109\/SmartNets61466.2024.10577645"},{"key":"382_CR68","unstructured":"Arjovsky M, Chintala S(2017) Bottou. wasserstein gan. arXiv preprint arXiv:1701.078757"},{"key":"382_CR69","unstructured":"Mescheder L, Nowozin S, Geiger A (2017) Adversarial variational bayes: Unifying variational autoencoders and generative adversarial networks. In: International Conference on Machine Learning, pp. 2391\u20132400"},{"key":"382_CR70","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2024.107896","volume":"132","author":"X Zhuang","year":"2024","unstructured":"Zhuang X, Li D, Wang Y, Li K (2024) Military target detection method based on efficientdet and generative adversarial network. Eng Appl Artif Intell 132:107896","journal-title":"Eng Appl Artif Intell"},{"key":"382_CR71","unstructured":"Sohl-Dickstein J, Weiss E, Maheswaranathan N, Ganguli S Deep (2015) unsupervised learning using nonequilibrium thermodynamics. In: International Conference on Machine Learning, pp. 2256\u20132265"},{"key":"382_CR72","first-page":"6840","volume":"33","author":"J Ho","year":"2020","unstructured":"Ho J, Jain A, Abbeel P (2020) Denoising diffusion probabilistic models. Adv Neural Inf Process Syst 33:6840\u20136851","journal-title":"Adv Neural Inf Process Syst"},{"key":"382_CR73","unstructured":"Nichol AQ, Dhariwal P (2021) Improved denoising diffusion probabilistic models. In: International Conference on Machine Learning, pp. 8162\u20138171"},{"key":"382_CR74","first-page":"8780","volume":"34","author":"P Dhariwal","year":"2021","unstructured":"Dhariwal P, Nichol A (2021) Diffusion models beat gans on image synthesis. Adv Neural Inf Process Syst 34:8780\u20138794","journal-title":"Adv Neural Inf Process Syst"},{"key":"382_CR75","unstructured":"Wang Y, Schiff Y, Gokaslan A, Pan W, Wang F, De\u00a0Sa C, Kuleshov V Infodiffusion: Representation learning using information maximizing diffusion models. In: International Conference on Machine Learning, pp. 36336\u201336354 (2023)"},{"key":"382_CR76","doi-asserted-by":"crossref","unstructured":"Yang X, Wang X (2023) Diffusion model as representation learner. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 18938\u201318949","DOI":"10.1109\/ICCV51070.2023.01736"},{"key":"382_CR77","doi-asserted-by":"publisher","first-page":"695","DOI":"10.1007\/s00371-013-0806-4","volume":"29","author":"R Song","year":"2013","unstructured":"Song R, Liu Y, Martin RR, Rosin PL (2013) 3d point of interest detection via spectral irregularity diffusion. Vis Comput 29:695\u2013705","journal-title":"Vis Comput"},{"key":"382_CR78","unstructured":"Krizhevsky A, Hinton G (2009) Learning multiple layers of features from tiny images[J\/OL]. Handbook of Systemic Autoimmune Diseases 1(4)"},{"issue":"3","key":"382_CR79","doi-asserted-by":"publisher","first-page":"249","DOI":"10.1097\/WAD.0b013e318142774e","volume":"21","author":"DL Beekly","year":"2007","unstructured":"Beekly DL, Ramos EM, Lee WW, Deitrich WD, Jacka ME, Wu J, Hubbard JL, Koepsell TD, Morris JC, Kukull WA et al (2007) The national alzheimer\u2019s coordinating center (nacc) database: the uniform data set. Alzheimer Disease & Associated Disorders 21(3):249\u2013258","journal-title":"Alzheimer Disease & Associated Disorders"},{"key":"382_CR80","doi-asserted-by":"crossref","unstructured":"Hariharan B, Girshick, R.: Low-shot visual recognition by shrinking and hallucinating features. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 3018\u20133027 (2017)","DOI":"10.1109\/ICCV.2017.328"},{"key":"382_CR81","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li L-J, Li K, Fei-Fei L (2009) Imagenet: A large-scale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248\u2013255","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"382_CR82","unstructured":"Wah C, Branson S, Welinder P, Perona P, Belongie S (2011) The caltech-ucsd birds-200-2011 dataset. California Institute of Technology"},{"key":"382_CR83","unstructured":"Wang T, Isola P (2020) Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In: International Conference on Machine Learning, pp. 9929\u20139939"},{"key":"382_CR84","doi-asserted-by":"crossref","unstructured":"Wu Z, Xiong Y, Yu SX, Lin D (2018) Unsupervised feature learning via non-parametric instance discrimination. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3733\u20133742","DOI":"10.1109\/CVPR.2018.00393"},{"key":"382_CR85","doi-asserted-by":"crossref","unstructured":"Ye M, Zhang X, Yuen PC, Chang S-F (2019) Unsupervised embedding learning via invariant and spreading instance feature. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 6210\u20136219","DOI":"10.1109\/CVPR.2019.00637"},{"key":"382_CR86","doi-asserted-by":"crossref","unstructured":"He K, Fan H, Wu Y, Xie S, Girshick R (2020) Momentum contrast for unsupervised visual representation learning. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 9729\u20139738","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"382_CR87","unstructured":"Chen T, Kornblith S, Norouzi M, Hinton G (2020) A simple framework for contrastive learning of visual representations. In: International Conference on Machine Learning, pp. 1597\u20131607"},{"key":"382_CR88","unstructured":"Oord Avd, Li Y, Vinyals O (2018) Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748"},{"key":"382_CR89","unstructured":"Chen X, Fan H, Girshick R, He K (2020) Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297"},{"key":"382_CR90","first-page":"22243","volume":"33","author":"T Chen","year":"2020","unstructured":"Chen T, Kornblith S, Swersky K, Norouzi M, Hinton GE (2020) Big self-supervised models are strong semi-supervised learners. Adv Neural Inf Process Syst 33:22243\u201322255","journal-title":"Adv Neural Inf Process Syst"},{"issue":"2","key":"382_CR91","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3463506","volume":"5","author":"H Haresamudram","year":"2021","unstructured":"Haresamudram H, Essa I, Pl\u00f6tz T (2021) Contrastive predictive coding for human activity recognition. Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies 5(2):1\u201326","journal-title":"Proceedings of the ACM on Interactive Mobile Wearable and Ubiquitous Technologies"},{"key":"382_CR92","unstructured":"Henaff O (2020) Data-efficient image recognition with contrastive predictive coding. In: International Conference on Machine Learning, pp. 4182\u20134192 . PMLR"},{"key":"382_CR93","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.108710","volume":"129","author":"J Yin","year":"2022","unstructured":"Yin J, Xie J, Ma Z, Guo J (2022) Mpccl: multiview predictive coding with contrastive learning for person re-identification. Pattern Recogn 129:108710","journal-title":"Pattern Recogn"},{"key":"382_CR94","doi-asserted-by":"crossref","unstructured":"Chao G, Jiang Y, Chu D (2024) Incomplete contrastive multi-view clustering with high-confidence guiding. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, pp. 11221\u201311229","DOI":"10.1609\/aaai.v38i10.29000"},{"key":"382_CR95","doi-asserted-by":"crossref","unstructured":"Huynh T, Kornblith S, Walter MR, Maire M, Khademi M (2022) Boosting contrastive self-supervised learning with false negative cancellation. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, pp. 2785\u20132795","DOI":"10.1109\/WACV51458.2022.00106"},{"key":"382_CR96","first-page":"9912","volume":"33","author":"M Caron","year":"2020","unstructured":"Caron M, Misra I, Mairal J, Goyal P, Bojanowski P, Joulin A (2020) Unsupervised learning of visual features by contrasting cluster assignments. Adv Neural Inf Process Syst 33:9912\u20139924","journal-title":"Adv Neural Inf Process Syst"},{"key":"382_CR97","doi-asserted-by":"crossref","unstructured":"Yin H, Vahdat A, Alvarez JM, Mallya A, Kautz J, Molchanov P (2022) A-vit: Adaptive tokens for efficient vision transformer. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 10809\u201310818","DOI":"10.1109\/CVPR52688.2022.01054"},{"issue":"4","key":"382_CR98","doi-asserted-by":"publisher","first-page":"1934","DOI":"10.1007\/s12083-024-01696-w","volume":"17","author":"SK Verma","year":"2024","unstructured":"Verma SK, Lokeshwaran K, Sahayaraj JM, Johnsana JA (2024) Energy efficient multi-objective cluster-based routing protocol for wsn using interval type-2 fuzzy logic modified dingo optimization. Peer-to-Peer Networking and Applications 17(4):1934\u20131962","journal-title":"Peer-to-Peer Networking and Applications"},{"key":"382_CR99","doi-asserted-by":"crossref","unstructured":"Van\u00a0Gansbeke W, Vandenhende S, Georgoulis S, Proesmans M, Van\u00a0Gool L (2020) Scan: Learning to classify images without labels. In: European Conference on Computer Vision, pp. 268\u2013285","DOI":"10.1007\/978-3-030-58607-2_16"},{"key":"382_CR100","doi-asserted-by":"crossref","unstructured":"Chen X, Xie S, He K (2021) An empirical study of training self-supervised vision transformers. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 9640\u20139649","DOI":"10.1109\/ICCV48922.2021.00950"},{"key":"382_CR101","doi-asserted-by":"crossref","unstructured":"Caron M, Touvron H, Misra I, J\u00e9gou H, Mairal J, Bojanowski P, Joulin A(2021) Emerging properties in self-supervised vision transformers. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 9650\u20139660","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"382_CR102","first-page":"21271","volume":"33","author":"J-B Grill","year":"2020","unstructured":"Grill J-B, Strub F, Altch\u00e9 F, Tallec C, Richemond P, Buchatskaya E, Doersch C, Avila Pires B, Guo Z, Gheshlaghi Azar M et al (2020) Bootstrap your own latent-a new approach to self-supervised learning. Adv Neural Inf Process Syst 33:21271\u201321284","journal-title":"Adv Neural Inf Process Syst"},{"key":"382_CR103","unstructured":"Fetterman A, Albrecht J (2020) Understanding self-supervised and contrastive learning with bootstrap your own latent (byol). Untitled AI, August"},{"key":"382_CR104","unstructured":"Tian Y, Yu L, Chen X, Ganguli S (2020) Understanding self-supervised learning with dual deep networks. arXiv preprint arXiv:2010.00578"},{"key":"382_CR105","unstructured":"Richemond PH, Grill J-B, Altch\u00e9 F, Tallec C, Strub F, Brock A, Smith S, De S, Pascanu R, Piot B, et al (2020) Byol works even without batch statistics. arXiv preprint arXiv:2010.10241"},{"key":"382_CR106","doi-asserted-by":"crossref","unstructured":"Chen X, He K (2021) Exploring simple siamese representation learning. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 15750\u201315758","DOI":"10.1109\/CVPR46437.2021.01549"},{"key":"382_CR107","unstructured":"Zbontar J, Jing L, Misra I, LeCun Y, Deny S (2021) Barlow twins: Self-supervised learning via redundancy reduction. In: International Conference on Machine Learning, pp. 12310\u201312320"},{"key":"382_CR108","doi-asserted-by":"crossref","unstructured":"Assran M, Caron M, Misra I, Bojanowski P, Bordes F, Vincent P, Joulin A, Rabbat M, Ballas N (2022) Masked siamese networks for label-efficient learning. In: European Conference on Computer Vision, pp. 456\u2013473 . Springer","DOI":"10.1007\/978-3-031-19821-2_26"},{"key":"382_CR109","unstructured":"Oquab M, Darcet T, Moutakanni T, Vo H, Szafraniec M, Khalidov V, Fernandez P, Haziza D, Massa F., El-Nouby A, et al (2023) Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193"},{"key":"382_CR110","doi-asserted-by":"publisher","first-page":"2348","DOI":"10.52202\/079017-0077","volume":"37","author":"S Mo","year":"2024","unstructured":"Mo S, Tong S (2024) Connecting joint-embedding predictive architecture with contrastive self-supervised learning. Adv Neural Inf Process Syst 37:2348\u20132377","journal-title":"Adv Neural Inf Process Syst"},{"key":"382_CR111","unstructured":"Mikolov T, Chen K, Corrado G, Dean J (2013) Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781"},{"key":"382_CR112","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. Advances in neural information processing systems 30"},{"key":"382_CR113","doi-asserted-by":"crossref","unstructured":"Devlin J, Chang M-W, Lee K, Toutanova K (2019) Bert: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (long and Short Papers), pp. 4171\u20134186","DOI":"10.18653\/v1\/N19-1423"},{"issue":"2","key":"382_CR114","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1109\/TPAMI.2018.2798607","volume":"41","author":"T Baltru\u0161aitis","year":"2018","unstructured":"Baltru\u0161aitis T, Ahuja C, Morency L-P (2018) Multimodal machine learning: a survey and taxonomy. IEEE Trans Pattern Anal Mach Intell 41(2):423\u2013443","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"382_CR115","doi-asserted-by":"crossref","unstructured":"Sun C, Myers A, Vondrick C, Murphy K, Schmid C (2019) Videobert: A joint model for video and language representation learning. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 7464\u20137473","DOI":"10.1109\/ICCV.2019.00756"},{"key":"382_CR116","unstructured":"Li LH, Yatskar M, Yin D, Hsieh C-J, Chang K-W (2019) Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557"},{"key":"382_CR117","doi-asserted-by":"crossref","unstructured":"Chen Y-C, Li L, Yu L, El\u00a0Kholy A, Ahmed F, Gan Z, Cheng Y, Liu J (2020) Uniter: Universal image-text representation learning. In: European Conference on Computer Vision, pp. 104\u2013120","DOI":"10.1007\/978-3-030-58577-8_7"},{"key":"382_CR118","unstructured":"Su W, Zhu X, Cao Y, Li B, Lu L, Wei F, Dai J (2019) Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530"},{"key":"382_CR119","doi-asserted-by":"crossref","unstructured":"Li G, Duan N, Fang Y, Gong M, Jiang D (2020) Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, pp. 11336\u201311344","DOI":"10.1609\/aaai.v34i07.6795"},{"key":"382_CR120","doi-asserted-by":"crossref","unstructured":"Alberti C, Ling J, Collins M, Reitter D (2019) Fusion of detected objects in text for visual question answering. arXiv preprint arXiv:1908.05054","DOI":"10.18653\/v1\/D19-1219"},{"key":"382_CR121","unstructured":"Huang Z, Zeng Z, Liu B, Fu D, Fu J (2020) Pixel-bert: Aligning image pixels with text by deep multi-modal transformers. arXiv preprint arXiv:2004.00849"},{"key":"382_CR122","doi-asserted-by":"crossref","unstructured":"Huang Z, Zeng Z, Huang Y, Liu B, Fu D, Fu J (2021) Seeing out of the box: End-to-end pre-training for vision-language representation learning. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 12976\u201312985","DOI":"10.1109\/CVPR46437.2021.01278"},{"key":"382_CR123","unstructured":"Kim W, Son B, Kim I (2021) Vilt: Vision-and-language transformer without convolution or region supervision. In: International Conference on Machine Learning, pp. 5583\u20135594"},{"key":"382_CR124","doi-asserted-by":"crossref","unstructured":"Li X, Yin X, Li C, Zhang P, Hu X, Zhang L, Wang L, Hu H, Dong L, Wei F (2020) et al.: Oscar: Object-semantics aligned pre-training for vision-language tasks. In: Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part XXX 16, pp. 121\u2013137","DOI":"10.1007\/978-3-030-58577-8_8"},{"key":"382_CR125","unstructured":"Zhang P, Li X, Hu X, Yang J, Zhang L, Wang L, Choi Y, Gao J (2021) Vinvl: Making visual representations matter in vision-language models. arXiv preprint arXiv:2101.00529 1(6), 8"},{"key":"382_CR126","unstructured":"Hu X, Yin X, Lin K, Wang L, Zhang L, Gao J, Liu Z (2020) Vivo: Surpassing human performance in novel object captioning with visual vocabulary pre-training. arXiv preprint arXiv:2009.13682 2(6), 17"},{"key":"382_CR127","first-page":"32897","volume":"35","author":"H Bao","year":"2022","unstructured":"Bao H, Wang W, Dong L, Liu Q, Mohammed OK, Aggarwal K, Som S, Piao S, Wei F (2022) Vlmo: unified vision-language pre-training with mixture-of-modality-experts. Adv Neural Inf Process Syst 35:32897\u201332912","journal-title":"Adv Neural Inf Process Syst"},{"key":"382_CR128","doi-asserted-by":"crossref","unstructured":"Wang W, Bao H, Dong L, Bjorck J, Peng Z, Liu Q, Aggarwal K, Mohammed OK, Singhal S, Som S, et al.: (2023) Image as a foreign language: Beit pretraining for vision and vision-language tasks. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 19175\u201319186","DOI":"10.1109\/CVPR52729.2023.01838"},{"key":"382_CR129","unstructured":"Lin XV, Shrivastava A, Luo L, Iyer S, Lewis M, Ghosh G, Zettlemoyer L, Aghajanyan A (2024) Moma: Efficient early-fusion pre-training with mixture of modality-aware experts. arXiv preprint arXiv:2407.21770"},{"key":"382_CR130","unstructured":"Lu J, Batra D, Parikh D, Lee S (2019) Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems 32"},{"key":"382_CR131","doi-asserted-by":"crossref","unstructured":"Tan H, Bansal M (2019) Lxmert: Learning cross-modality encoder representations from transformers. arXiv preprint arXiv:1908.07490","DOI":"10.18653\/v1\/D19-1514"},{"key":"382_CR132","doi-asserted-by":"crossref","unstructured":"Lu J, Goswami V, Rohrbach M, Parikh D, Lee S (2020) 12-in-1: Multi-task vision and language representation learning. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 10437\u201310446","DOI":"10.1109\/CVPR42600.2020.01045"},{"key":"382_CR133","unstructured":"Sun Y, Wang S, Li Y, Feng S, Chen X, Zhang H, Tian X, Zhu D, Tian H, Wu H (2019) Ernie: Enhanced representation through knowledge integration. arXiv preprint arXiv:1904.09223"},{"key":"382_CR134","doi-asserted-by":"crossref","unstructured":"Yu F, Tang J, Yin W, Sun Y, Tian H, Wu H, Wang H (2021) Ernie-vil: Knowledge enhanced vision-language representations through scene graphs. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 3208\u20133216","DOI":"10.1609\/aaai.v35i4.16431"},{"key":"382_CR135","unstructured":"Li J, Li D, Xiong C, Hoi S (2022) Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In: International Conference on Machine Learning, pp. 12888\u201312900 . PMLR"},{"key":"382_CR136","unstructured":"Li J, Li D, Savarese S, Hoi S (2023) Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In: International Conference on Machine Learning, pp. 19730\u201319742 . PMLR"},{"key":"382_CR137","unstructured":"Sun Q, Wang J, Yu Q, Cui Y, Zhang F, Zhang X, Wang X (2024) Eva-clip-18b: Scaling clip to 18 billion parameters. arXiv preprint arXiv:2402.04252"},{"key":"382_CR138","unstructured":"Li C, Yan M, Xu H, Luo F, Wang W, Bi B, Huang S (2021) Semvlp: Vision-language pre-training by aligning semantics at multiple levels. arXiv preprint arXiv:2103.07829"},{"key":"382_CR139","unstructured":"Yu J, Wang Z, Vasudevan V, Yeung L, Seyedhosseini M, Wu Y (2022) Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917"},{"key":"382_CR140","doi-asserted-by":"crossref","unstructured":"Chen X, Djolonga J, Padlewski P, Mustafa B, Changpinyo S, Wu J, Ruiz CR, Goodman S, Wang X, Tay Y, et al (2023) Pali-x: On scaling up a multilingual vision and language model. arXiv preprint arXiv:2305.18565","DOI":"10.1109\/CVPR52733.2024.01368"},{"key":"382_CR141","unstructured":"Chen Z, Wang W, Cao Y, Liu Y, Gao Z, Cui E, Zhu J, Ye S, Tian H, Liu Z, et al (2024) Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271"},{"key":"382_CR142","doi-asserted-by":"crossref","unstructured":"Lee K-H, Chen X, Hua G, Hu H, He X (2018) Stacked cross attention for image-text matching. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 201\u2013216","DOI":"10.1007\/978-3-030-01225-0_13"},{"key":"382_CR143","unstructured":"Faghri F, Fleet DJ, Kiros JR, Fidler S (2017) Vse++: Improving visual-semantic embeddings with hard negatives. arXiv preprint arXiv:1707.05612"},{"key":"382_CR144","unstructured":"Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, et al.: (2021) Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, pp. 8748\u20138763"},{"key":"382_CR145","unstructured":"Jia C, Yang Y, Xia Y, Chen Y-T, Parekh Z, Pham H, Le Q, Sung Y-H, Li Z, Duerig T (2021) Scaling up visual and vision-language representation learning with noisy text supervision. In: International Conference on Machine Learning, pp. 4904\u20134916"},{"key":"382_CR146","doi-asserted-by":"publisher","first-page":"424","DOI":"10.1016\/j.inffus.2022.09.025","volume":"91","author":"A Gandhi","year":"2023","unstructured":"Gandhi A, Adhvaryu K, Poria S, Cambria E, Hussain A (2023) Multimodal sentiment analysis: a systematic review of history, datasets, multimodal fusion methods, applications, challenges and future directions. Information Fusion 91:424\u2013444","journal-title":"Information Fusion"},{"key":"382_CR147","doi-asserted-by":"crossref","unstructured":"Farhadi A, Endres I, Hoiem D, Forsyth D (2009) Describing objects by their attributes. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1778\u20131785","DOI":"10.1109\/CVPR.2009.5206772"},{"key":"382_CR148","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll\u00e1r P, Zitnick CL (2014) Microsoft coco: Common objects in context. In: Computer vision\u2013ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pp. 740\u2013755","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"382_CR149","doi-asserted-by":"crossref","unstructured":"Plummer BA, Wang L, Cervantes CM, Caicedo JC, Hockenmaier J, Lazebnik S (2015) Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 2641\u20132649","DOI":"10.1109\/ICCV.2015.303"},{"key":"382_CR150","doi-asserted-by":"crossref","unstructured":"Antol S, Agrawal A, Lu J, Mitchell M, Batra D, Zitnick CL, Parikh D (2015) Vqa: Visual question answering. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 2425\u20132433","DOI":"10.1109\/ICCV.2015.279"},{"key":"382_CR151","doi-asserted-by":"crossref","unstructured":"Goyal Y, Khot T, Summers-Stay D, Batra D, Parikh D (2017) Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6904\u20136913","DOI":"10.1109\/CVPR.2017.670"},{"key":"382_CR152","doi-asserted-by":"crossref","unstructured":"Hudson DA, Manning CD (2019) Gqa: A new dataset for real-world visual reasoning and compositional question answering. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 6700\u20136709","DOI":"10.1109\/CVPR.2019.00686"},{"key":"382_CR153","unstructured":"Ren M, Kiros R, Zemel R (2015) Exploring models and data for image question answering. Advances in neural information processing systems 28"},{"key":"382_CR154","doi-asserted-by":"crossref","unstructured":"Johnson J, Hariharan B, Van Der\u00a0Maaten L, Fei-Fei L, Lawrence\u00a0Zitnick C, Girshick R (2017) Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2901\u20132910","DOI":"10.1109\/CVPR.2017.215"},{"key":"382_CR155","doi-asserted-by":"crossref","unstructured":"Chua T-S, Tang J, Hong R, Li H, Luo Z, Zheng Y (2009) Nus-wide: a real-world web image database from national university of singapore. In: Proceedings of the ACM International Conference on Image and Video Retrieval, pp. 1\u20139","DOI":"10.1145\/1646396.1646452"},{"key":"382_CR156","doi-asserted-by":"crossref","unstructured":"Sharma P, Ding N, Goodman S, Soricut R (2018) Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2556\u20132565","DOI":"10.18653\/v1\/P18-1238"},{"key":"382_CR157","unstructured":"Alayrac J-B, Donahue J, Luc P, Miech A, Barr I, Hasson Y, Lenc K, Mensch A, Millican K, Reynolds M et al (2022) Flamingo: a visual language model for few-shot learning. Adv Neural Inf Process Syst 35:23716\u201323736"},{"key":"382_CR158","doi-asserted-by":"crossref","unstructured":"Sharma P, Ding N, Goodman S, Soricut R (2018) Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, vol. 1, pp. 2556\u20132565","DOI":"10.18653\/v1\/P18-1238"}],"container-title":["International Journal of Multimedia Information Retrieval"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13735-025-00382-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s13735-025-00382-8","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13735-025-00382-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,20]],"date-time":"2025-12-20T07:04:18Z","timestamp":1766214258000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s13735-025-00382-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,30]]},"references-count":158,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,12]]}},"alternative-id":["382"],"URL":"https:\/\/doi.org\/10.1007\/s13735-025-00382-8","relation":{},"ISSN":["2192-6611","2192-662X"],"issn-type":[{"value":"2192-6611","type":"print"},{"value":"2192-662X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,30]]},"assertion":[{"value":"17 May 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 September 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 September 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 September 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of Interest"}}],"article-number":"32"}}