{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T05:27:21Z","timestamp":1781587641506,"version":"3.54.5"},"reference-count":71,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2025,6,1]],"date-time":"2025-06-01T00:00:00Z","timestamp":1748736000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,6,2]],"date-time":"2025-06-02T00:00:00Z","timestamp":1748822400000},"content-version":"vor","delay-in-days":1,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Johannes Kepler University Linz"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Multimed Info Retr"],"published-print":{"date-parts":[[2025,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Multimodal learning has demonstrated remarkable performance improvements over unimodal architectures. However, multimodal learning methods often exhibit deteriorated performances if one or more modalities are missing. This may be attributed to the commonly used multi-branch design containing modality-specific components, making such approaches reliant on the availability of a complete set of modalities. In this work, we propose a robust multimodal learning framework, , that adapts a common-space visual learning network to align all input modalities. To enable this, we present the unification of input modalities into one format by encoding any non-visual modality into visual representations thus making it robust to missing modalities. Extensive experiments are performed on multimodal classification task using four textual-visual (Hateful Memes, UPMC Food-101, MM-IMDb, and Ferramenta) and two audio-visual (avMNIST, VoxCeleb) datasets.  not only achieves superior performance when all modalities are present at train\/test time but also demonstrates notable resilience in the case of missing modalities.<\/jats:p>","DOI":"10.1007\/s13735-025-00370-y","type":"journal-article","created":{"date-parts":[[2025,6,2]],"date-time":"2025-06-02T07:25:19Z","timestamp":1748849119000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Chameleon: A Multimodal Learning Framework Robust to Missing Modalities"],"prefix":"10.1007","volume":"14","author":[{"given":"Muhammad Irzam","family":"Liaqat","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shah","family":"Nawaz","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Muhammad Zaigham","family":"Zaheer","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Muhammad Saad","family":"Saeed","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hassan","family":"Sajjad","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tom","family":"De Schepper","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Karthik","family":"Nandakumar","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Muhammad Haris","family":"Khan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ignazio","family":"Gallo","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Markus","family":"Schedl","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,6,2]]},"reference":[{"key":"370_CR1","doi-asserted-by":"crossref","unstructured":"Kiela D, Grave E, Joulin A, Mikolov T (2018). Efficient large-scale multi-modal classification. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, 5198\u20135204","DOI":"10.1609\/aaai.v32i1.11945"},{"key":"370_CR2","first-page":"2611","volume":"33","author":"D Kiela","year":"2020","unstructured":"Kiela D, Firooz H, Mohan A, Goswami V, Singh A, Ringshia P, Testuggine D (2020) The hateful memes challenge: Detecting hate speech in multimodal memes. Advances in Neural Information Processing Systems 33:2611\u20132624","journal-title":"Advances in Neural Information Processing Systems"},{"key":"370_CR3","doi-asserted-by":"crossref","unstructured":"Wang L, Li Y, Lazebnik S (2016). Learning deep structure-preserving image-text embeddings. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5005\u20135013","DOI":"10.1109\/CVPR.2016.541"},{"key":"370_CR4","doi-asserted-by":"crossref","unstructured":"Nagrani A, Albanie S, Zisserman A (2018). Seeing voices and hearing faces: Cross-modal biometric matching. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8427\u20138436","DOI":"10.1109\/CVPR.2018.00879"},{"key":"370_CR5","doi-asserted-by":"publisher","unstructured":"Moon S, Neves L, Carvalho V (2018). Multimodal named entity recognition for short social media posts. In: Walker, M., Ji, H., Stent, A. (eds.) Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 852\u2013860. Association for Computational Linguistics, New Orleans, Louisiana . https:\/\/doi.org\/10.18653\/v1\/N18-1078 . https:\/\/aclanthology.org\/N18-1078\/","DOI":"10.18653\/v1\/N18-1078"},{"key":"370_CR6","doi-asserted-by":"crossref","unstructured":"Arshad O, Gallo I, Nawaz S, Calefati A (2019). Aiding intra-text representations with visual context for multimodal named entity recognition. In: 2019 International Conference on Document Analysis and Recognition (ICDAR), pp. 337\u2013342","DOI":"10.1109\/ICDAR.2019.00061"},{"key":"370_CR7","doi-asserted-by":"crossref","unstructured":"Anderson P, He X, Buehler C, Teney D, Johnson M, Gould S, Zhang L(2018). Bottom-up and top-down attention for image captioning and visual question answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6077\u20136086","DOI":"10.1109\/CVPR.2018.00636"},{"key":"370_CR8","doi-asserted-by":"crossref","unstructured":"Fukui A, Park D.H, Yang D, Rohrbach A, Darrell T, Rohrbach M (2016). Multimodal compact bilinear pooling for visual question answering and visual grounding. arXiv:1606.01847","DOI":"10.18653\/v1\/D16-1044"},{"key":"370_CR9","doi-asserted-by":"crossref","unstructured":"Vinyals O, Toshev A, Bengio S, Erhan D (2015). Show and tell: A neural image caption generator. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3156\u20133164","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"370_CR10","doi-asserted-by":"crossref","unstructured":"Kiela D, Bottou L (2014). Learning image embeddings using convolutional neural networks for improved multi-modal semantics. In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 36\u201345","DOI":"10.3115\/v1\/D14-1005"},{"issue":"12","key":"370_CR11","doi-asserted-by":"publisher","first-page":"5827","DOI":"10.1109\/JBHI.2023.3319361","volume":"27","author":"Y Tian","year":"2023","unstructured":"Tian Y, Jian G, Wang J, Chen H, Pan L, Xu Z, Li J, Wang R (2023) A revised approach to orthodontic treatment monitoring from oralscan video. IEEE Journal of Biomedical and Health Informatics 27(12):5827\u20135836","journal-title":"IEEE Journal of Biomedical and Health Informatics"},{"key":"370_CR12","doi-asserted-by":"crossref","unstructured":"Specia L, Frank S, Sima\u2019An K, Elliott D (2016). A shared task on multimodal machine translation and crosslingual image description. In: Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, pp. 543\u2013553","DOI":"10.18653\/v1\/W16-2346"},{"key":"370_CR13","doi-asserted-by":"crossref","unstructured":"Elliott D, Frank S, Sima\u2019an K, Specia L (2016). Multi30k: Multilingual english-german image descriptions. arXiv:1605.00459","DOI":"10.18653\/v1\/W16-3210"},{"key":"370_CR14","doi-asserted-by":"crossref","unstructured":"Ma M, Ren J, Zhao L, Testuggine D, Peng X (2022). Are multimodal transformers robust to missing modality? In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 18177\u201318186","DOI":"10.1109\/CVPR52688.2022.01764"},{"key":"370_CR15","doi-asserted-by":"crossref","unstructured":"Ma M, Ren J, Zhao L, Tulyakov S, Wu C, Peng X (2021). Smil: Multimodal learning with severely missing modality. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 2302\u20132310","DOI":"10.1609\/aaai.v35i3.16330"},{"key":"370_CR16","doi-asserted-by":"crossref","unstructured":"Lee Y.-L, Tsai Y.-H, Chiu W.-C, Lee C.-Y (2023). Multimodal prompting with missing modalities for visual recognition. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 14943\u201314952","DOI":"10.1109\/CVPR52729.2023.01435"},{"key":"370_CR17","unstructured":"Kim W, Son B, Kim I (2021). Vilt: Vision-and-language transformer without convolution or region supervision. In: International Conference on Machine Learning, pp. 5583\u20135594"},{"key":"370_CR18","unstructured":"Wang X, Kumar D, Thome N, Cord M, Precioso F (2015). Recipe recognition with large multimodal food dataset. In: 2015 IEEE International Conference on Multimedia & Expo Workshops (ICMEW), pp. 1\u20136"},{"key":"370_CR19","unstructured":"Arevalo J, Solorio T, Montes-y-G\u00f3mez M, Gonz\u00e1lez F.A (2017). Gated multimodal units for information fusion. arXiv:1702.01992"},{"key":"370_CR20","doi-asserted-by":"crossref","unstructured":"Gallo I, Calefati A, Nawaz S (2017) Multimodal classification fusion in real-world scenarios. In: 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), vol. 5, pp. 36\u201341","DOI":"10.1109\/ICDAR.2017.326"},{"key":"370_CR21","doi-asserted-by":"crossref","unstructured":"Vielzeuf V, Lechervy A, Pateux S, Jurie F (2018). Centralnet: a multilayer approach for multimodal fusion. In: Proceedings of the European Conference on Computer Vision (ECCV) Workshops, pp. 0\u20130","DOI":"10.1007\/978-3-030-11024-6_44"},{"key":"370_CR22","doi-asserted-by":"crossref","unstructured":"Nagrani A, Chung J.S, Zisserman A (2017). Voxceleb: a large-scale speaker identification dataset. In: INTERSPEECH","DOI":"10.21437\/Interspeech.2017-950"},{"issue":"2","key":"370_CR23","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1109\/TPAMI.2018.2798607","volume":"41","author":"T Baltru\u0161aitis","year":"2018","unstructured":"Baltru\u0161aitis T, Ahuja C, Morency L-P (2018) Multimodal machine learning: A survey and taxonomy. IEEE transactions on pattern analysis and machine intelligence 41(2):423\u2013443","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"issue":"10","key":"370_CR24","doi-asserted-by":"publisher","first-page":"12113","DOI":"10.1109\/TPAMI.2023.3275156","volume":"45","author":"P Xu","year":"2023","unstructured":"Xu P, Zhu X, Clifton DA (2023) Multimodal learning with transformers: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 45(10):12113\u201312132","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"370_CR25","doi-asserted-by":"crossref","unstructured":"Nagrani A, Albanie S, Zisserman A (2018). Learnable pins: Cross-modal embeddings for person identity. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 71\u201388","DOI":"10.1007\/978-3-030-01261-8_5"},{"key":"370_CR26","doi-asserted-by":"crossref","unstructured":"Saeed M.S, Khan M.H, Nawaz S, Yousaf M.H, Del\u00a0Bue A (2022). Fusion and orthogonal projection for improved face-voice association. In: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 7057\u20137061","DOI":"10.1109\/ICASSP43922.2022.9747704"},{"key":"370_CR27","doi-asserted-by":"crossref","unstructured":"Kim C, Shin H.V, Oh T.-H, Kaspar A, Elgharib M, Matusik W (2018). On learning associations of faces and voices. In: Asian Conference on Computer Vision, pp. 276\u2013292 . Springer","DOI":"10.1007\/978-3-030-20873-8_18"},{"key":"370_CR28","unstructured":"Lu J, Batra D, Parikh D, Lee S (2019). Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems 32"},{"key":"370_CR29","unstructured":"Radford A, Kim J.W, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, et al (2021). Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, pp. 8748\u20138763 . PMLR"},{"key":"370_CR30","doi-asserted-by":"crossref","unstructured":"He X, Peng Y (2017). Fine-grained image classification via combining vision and language. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5994\u20136002","DOI":"10.1109\/CVPR.2017.775"},{"key":"370_CR31","doi-asserted-by":"crossref","unstructured":"Yang F, Peng X, Ghosh G, Shilon R, Ma H, Moore E, Predovic G (2019). Exploring deep multimodal fusion of text and photo for hate speech classification. In: Proceedings of the Third Workshop on Abusive Language Online, pp. 11\u201318","DOI":"10.18653\/v1\/W19-3502"},{"issue":"13","key":"370_CR32","doi-asserted-by":"publisher","first-page":"2312","DOI":"10.3390\/math10132312","volume":"10","author":"X Zhang","year":"2022","unstructured":"Zhang X, Song Q, Liu G (2022) Multimodal image aesthetic prediction with missing modality. Mathematics 10(13):2312","journal-title":"Mathematics"},{"key":"370_CR33","doi-asserted-by":"crossref","unstructured":"Suo Q, Zhong W, Ma F, Yuan Y, Gao J, Zhang A (2019). Metric learning on healthcare data with incomplete modalities. In: IJCAI, pp. 3534\u20133540","DOI":"10.24963\/ijcai.2019\/490"},{"key":"370_CR34","unstructured":"Li M, Yang D, Liu Y, Wang S, Chen J, Wang S, Wei J, Jiang Y, Xu Q, Hou X, et al (2024). Toward robust incomplete multimodal sentiment analysis via hierarchical representation learning. arXiv:2411.02793"},{"key":"370_CR35","doi-asserted-by":"crossref","unstructured":"Zhang C, Chu X, Ma L, Zhu Y, Wang Y, Wang J, Zhao J (2022) M3care: Learning with missing modalities in multimodal healthcare data. In: Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 2418\u20132428","DOI":"10.1145\/3534678.3539388"},{"key":"370_CR36","doi-asserted-by":"crossref","unstructured":"Wang N, Cao H, Zhao J, Chen R, Yan D, Zhang J (2022). M2r2: Missing-modality robust emotion recognition framework with iterative data augmentation. IEEE Transactions on Artificial Intelligence","DOI":"10.1109\/TAI.2022.3201809"},{"key":"370_CR37","doi-asserted-by":"crossref","unstructured":"Wang H, Chen Y, Ma C, Avery J, Hull L, Carneiro G (2023). Multi-modal learning with missing modality via shared-specific feature modelling. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 15878\u201315887","DOI":"10.1109\/CVPR52729.2023.01524"},{"key":"370_CR38","doi-asserted-by":"crossref","unstructured":"Lan G, Du Y, Yang Z (2024) Robust multimodal representation under uncertain missing modalities. ACM Transactions on Multimedia Computing, Communications and Applications","DOI":"10.1145\/3702003"},{"key":"370_CR39","doi-asserted-by":"crossref","unstructured":"Gallo I, Nawaz S, Calefati A (2017). Semantic text encoding for text classification using convolutional neural networks. In: 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), vol. 5, pp. 16\u201321","DOI":"10.1109\/ICDAR.2017.323"},{"key":"370_CR40","doi-asserted-by":"crossref","unstructured":"Salesky E, Etter D, Post M (2021). Robust open-vocabulary translation from visual text representations. arXiv:2104.08211","DOI":"10.18653\/v1\/2021.emnlp-main.576"},{"key":"370_CR41","unstructured":"Rust P, Lotz J.F, Bugliarello E, Salesky E, Lhoneux M, Elliott D (2022). Language modelling with pixels. arXiv:2207.06991"},{"key":"370_CR42","doi-asserted-by":"crossref","unstructured":"Tschannen M, Mustafa B, Houlsby N (2023). Clippo: Image-and-language understanding from pixels only. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 11006\u201311017","DOI":"10.1109\/CVPR52729.2023.01059"},{"key":"370_CR43","doi-asserted-by":"crossref","unstructured":"Xie W, Nagrani A, Chung J.S, Zisserman A (2019). Utterance-level aggregation for speaker recognition in the wild. In: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5791\u20135795","DOI":"10.1109\/ICASSP.2019.8683120"},{"key":"370_CR44","doi-asserted-by":"crossref","unstructured":"Nawaz S, Janjua M.K, Gallo I, Mahmood A, Calefati A (2019). Deep latent space learning for cross-modal mapping of audio and visual signals. In: 2019 Digital Image Computing: Techniques and Applications (DICTA), pp. 1\u20137","DOI":"10.1109\/DICTA47822.2019.8945863"},{"key":"370_CR45","doi-asserted-by":"crossref","unstructured":"Selvaraju R.R, Cogswell M, Das A, Vedantam R, Parikh D, Batra D (2017). Grad-cam: Visual explanations from deep networks via gradient-based localization. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 618\u2013626","DOI":"10.1109\/ICCV.2017.74"},{"key":"370_CR46","unstructured":"Kiela D, Firooz H, Mohan A, Goswami V, Singh A, Fitzpatrick C.A, Bull P, Lipstein G, Nelli T, Zhu R, et al (2021). The hateful memes challenge: competition report. In: NeurIPS 2020 Competition and Demonstration Track, pp. 344\u2013360"},{"key":"370_CR47","unstructured":"Li X, Li J (2023). Angle-optimized text embeddings. arXiv:2309.12871"},{"key":"370_CR48","doi-asserted-by":"crossref","unstructured":"Desplanques B, Thienpondt J, Demuynck K (2020). Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification. arXiv:2005.07143","DOI":"10.21437\/Interspeech.2020-2650"},{"issue":"1","key":"370_CR49","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/LSENS.2018.2880790","volume":"3","author":"S Nawaz","year":"2018","unstructured":"Nawaz S, Calefati A, Janjua MK, Anwaar MU, Gallo I (2018) Learning fused representations for large-scale multimodal classification. IEEE Sensors Letters 3(1):1\u20134","journal-title":"IEEE Sensors Letters"},{"key":"370_CR50","unstructured":"Kiela D, Bhooshan S, Firooz H, Perez E, Testuggine D (2019). Supervised multimodal bitransformers for classifying images and text. arXiv preprint arXiv:1909.02950"},{"key":"370_CR51","doi-asserted-by":"crossref","unstructured":"Gallo I, Ria G, Landro N, La\u00a0Grassa R (2020). Image and text fusion for upmc food-101 using bert and cnns. In: 2020 35th International Conference on Image and Vision Computing New Zealand (IVCNZ), pp. 1\u20136 . IEEE","DOI":"10.1109\/IVCNZ51579.2020.9290732"},{"key":"370_CR52","unstructured":"Li L, Yatskar M, Yin D, Hsieh C, Chang K (2019). A simple and performant baseline for vision and language. arXiv:1908.03557"},{"key":"370_CR53","doi-asserted-by":"publisher","unstructured":"Fukui A, Park D.H, Yang D, Rohrbach A, Darrell T, Rohrbach M (2016). Multimodal compact bilinear pooling for visual question answering and visual grounding. In: Su, J., Duh, K., Carreras, X. (eds.) Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 457\u2013468. Association for Computational Linguistics, Austin, Texas . https:\/\/doi.org\/10.18653\/v1\/D16-1044 . https:\/\/aclanthology.org\/D16-1044","DOI":"10.18653\/v1\/D16-1044"},{"key":"370_CR54","doi-asserted-by":"crossref","unstructured":"P\u00e9rez-R\u00faa J.-M, Vielzeuf V, Pateux S, Baccouche M, Jurie F (2019). Mfas: Multimodal fusion architecture search. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 6966\u20136975","DOI":"10.1109\/CVPR.2019.00713"},{"key":"370_CR55","doi-asserted-by":"crossref","unstructured":"Gallo I, Calefati A, Nawaz S, Janjua M.K (2018). Image and encoded text fusion for multi-modal classification. In: 2018 Digital Image Computing: Techniques and Applications (DICTA), pp. 1\u20137","DOI":"10.1109\/DICTA.2018.8615789"},{"key":"370_CR56","doi-asserted-by":"crossref","unstructured":"Yue T, Li Y, Qin J, Hu Z (2023). Multi-modal hierarchical fusion network for fine-grained paper classification. Multimedia Tools and Applications, 1\u201317","DOI":"10.1007\/s11042-023-16626-w"},{"issue":"8","key":"370_CR57","doi-asserted-by":"publisher","first-page":"1692","DOI":"10.1109\/TPAMI.2015.2461544","volume":"38","author":"N Neverova","year":"2015","unstructured":"Neverova N, Wolf C, Taylor G, Nebout F (2015) Moddrop: adaptive multi-modal gesture recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 38(8):1692\u20131706","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"370_CR58","doi-asserted-by":"crossref","unstructured":"Saeed M.S, Nawaz S, Khan M.H, Zaheer M.Z, Nandakumar K, Yousaf M.H, Mahmood A (2023). Single-branch network for multimodal training. In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1\u20135","DOI":"10.1109\/ICASSP49357.2023.10097207"},{"key":"370_CR59","unstructured":"Liu H, Li C, Wu Q, Lee Y.J (2023). Visual instruction tuning. In: NeurIPS"},{"key":"370_CR60","doi-asserted-by":"crossref","unstructured":"Lee H.-C, Lin C.-Y, Hsu P.-C, Hsu W.H (2019). Audio feature generation for missing modality problem in video action recognition. In: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 3956\u20133960","DOI":"10.1109\/ICASSP.2019.8682513"},{"key":"370_CR61","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016). Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"370_CR62","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, et al (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929"},{"key":"370_CR63","doi-asserted-by":"crossref","unstructured":"Marouf I.E, Tartaglione E, Lathuili\u00e8re S (2024). Mini but mighty: Finetuning vits with mini adapters. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, pp. 1732\u20131741","DOI":"10.1109\/WACV57701.2024.00175"},{"key":"370_CR64","unstructured":"Wen Y, Ismail M.A, Liu W, Raj B, Singh R (2019). Disjoint mapping network for cross-modal matching of voices and faces. In: 7th International Conference on Learning Representations, ICLR 2019, USA, May 6-9, 2019"},{"key":"370_CR65","doi-asserted-by":"crossref","unstructured":"Nawaz S, Saeed M.S, Morerio P, Mahmood A, Gallo I, Yousaf M.H, Del\u00a0Bue A (2021). Cross-modal speaker verification and recognition: A multilingual perspective. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 1682\u20131691","DOI":"10.1109\/CVPRW53098.2021.00184"},{"key":"370_CR66","doi-asserted-by":"crossref","unstructured":"Sar\u0131 L, Singh K, Zhou J, Torresani L, Singhal N, Saraf Y (2021). A multi-view approach to audio-visual speaker verification. In: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 6194\u20136198 . IEEE","DOI":"10.1109\/ICASSP39728.2021.9414260"},{"key":"370_CR67","doi-asserted-by":"crossref","unstructured":"Zheng A, Hu M, Jiang B, Huang Y, Yan Y, Luo B (2021) Adversarial-metric learning for audio-visual cross-modal matching. IEEE Transactions on Multimedia 24:338\u2013351","DOI":"10.1109\/TMM.2021.3050089"},{"key":"370_CR68","doi-asserted-by":"publisher","first-page":"1763","DOI":"10.1109\/TMM.2021.3071243","volume":"24","author":"H Ning","year":"2021","unstructured":"Ning H, Zheng X, Lu X, Yuan Y (2021) Disentangled representation learning for cross-modal biometric matching. IEEE Transactions on Multimedia 24:1763\u20131774","journal-title":"IEEE Transactions on Multimedia"},{"key":"370_CR69","unstructured":"Song K, Tan X, Qin T, Lu J, Liu T-Y (2020) Mpnet: Masked and permuted pre-training for language understanding. Advances in neural information processing systems 33:16857\u201316867"},{"key":"370_CR70","doi-asserted-by":"publisher","unstructured":"Devlin J, Chang M.-W, Lee K, Toutanova K (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In: Burstein, J., Doran, C., Solorio, T. (eds.) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171\u20134186. Association for Computational Linguistics, Minneapolis, Minnesota . https:\/\/doi.org\/10.18653\/v1\/N19-1423https:\/\/aclanthology.org\/N19-1423","DOI":"10.18653\/v1\/N19-1423"},{"key":"370_CR71","unstructured":"Mikolov T, Sutskever I, Chen K, Corrado G.S, Dean J (2013). Distributed representations of words and phrases and their compositionality. In: Advances in Neural Information Processing Systems, pp. 3111\u20133119"}],"container-title":["International Journal of Multimedia Information Retrieval"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13735-025-00370-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s13735-025-00370-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s13735-025-00370-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,10]],"date-time":"2025-06-10T13:41:15Z","timestamp":1749562875000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s13735-025-00370-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6]]},"references-count":71,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,6]]}},"alternative-id":["370"],"URL":"https:\/\/doi.org\/10.1007\/s13735-025-00370-y","relation":{},"ISSN":["2192-6611","2192-662X"],"issn-type":[{"value":"2192-6611","type":"print"},{"value":"2192-662X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6]]},"assertion":[{"value":"24 January 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 April 2025","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 April 2025","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 June 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"21"}}