{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T17:01:49Z","timestamp":1781715709173,"version":"3.54.5"},"reference-count":9,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGKDD Explor. Newsl."],"published-print":{"date-parts":[[2026,6,17]]},"abstract":"<jats:p>Multimodal deep learning has achieved remarkable progress by leveraging complementary information across heterogeneous data sources such as texts, images, audios, and structured signals. While increasingly powerful encoders and fusion mechanisms have improved predictive performance, the reliability of multimodal systems remains a critical challenge. In particular, modality disagreement, distribution shifts, and noisy inputs can lead to overconfident yet incorrect predictions. Distinct from existing surveys on uncertainty in deep learning [45] or on multimodal learning [8; 135], this survey jointly covers three aspects: (i) the structural foundations of multimodal classification examined through the lens of the un- certainty challenges each design choice introduces; (ii) un- certainty quantification in both unimodal and multimodal settings; and (iii) set-valued classification as a decision-level strategy for cautious multimodal prediction. We first review foundational aspects of multimodal representation learning and fusion strategies, highlighting their structural limitations in modeling inter-modal dependence and the uncertainty challenges each stage introduces. We then examine uncertainty quantification methods in deep learning, including both probabilistic and evidence-theoretic approaches, and analyze how these techniques extend to multimodal settings. Special attention is given to conflict-aware fusion mechanisms and to decision-level strategies such as set- valued classification, which enable more cautious and informative predictions. Beyond reviewing existing methods, we identify key open challenges, including the modeling of partial dependence be- tween modalities, the need for systematic benchmarking of multimodal uncertainty, and the integration of uncertainty into decision-making pipelines. Finally, we discuss how these reliability challenges extend to emerging multimodal agentic systems. By synthesizing advances across multimodal learning and uncertainty modeling, this survey aims to provide a unified perspective and to outline recent research directions toward more reliable multimodal AI systems.<\/jats:p>","DOI":"10.1145\/3820356.3820360","type":"journal-article","created":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T15:58:43Z","timestamp":1781711923000},"page":"41-62","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Classification with Uncertainty-Aware Multimodal Deep Learning: A Survey"],"prefix":"10.1145","volume":"28","author":[{"given":"Grigor","family":"Bezirganyan","sequence":"first","affiliation":[{"name":"Aix Marseille Univ CNRS, LIS Marseille, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sana","family":"Sellami","sequence":"additional","affiliation":[{"name":"Aix Marseille Univ CNRS, LIS Marseille, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Laure","family":"Berti-\u00c9quille","sequence":"additional","affiliation":[{"name":"IRD, ESPACE-DEV Montpellier, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"S\u00e9bastien","family":"Fournier","sequence":"additional","affiliation":[{"name":"Aix Marseille Univ CNRS, LIS Marseille, France"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,17]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2021.05.008"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2022.09.023"},{"key":"e_1_2_1_3_1","volume-title":"Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716-23736","author":"Alayrac J.-B.","year":"2022","unstructured":"J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716-23736, 2022."},{"key":"e_1_2_1_4_1","volume-title":"PMLR","author":"Andrew G.","year":"2013","unstructured":"G. Andrew, R. Arora, J. Bilmes, and K. Livescu. Deep canonical correlation analysis. In International confer-ence on machine learning, pages 1247-1255. PMLR, 2013."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-010-0182-0"},{"key":"e_1_2_1_6_1","first-page":"427","volume-title":"Joint European confer-ence on machine learning and knowledge discovery in databases","author":"Audebert N.","year":"2019","unstructured":"N. Audebert, C. Herold, K. Slimani, and C. Vidal. Multimodal deep networks for text and image-based document classification. In Joint European confer-ence on machine learning and knowledge discovery in databases, pages 427-443. Springer, 2019."},{"key":"e_1_2_1_7_1","volume-title":"Medical Imaging with Deep Learning","author":"Ayhan M. S.","year":"2018","unstructured":"M. S. Ayhan and P. Berens. Test-time data augmen-tation for estimation of heteroscedastic aleatoric un-certainty in deep neural networks. In Medical Imaging with Deep Learning, 2018."},{"key":"e_1_2_1_8_1","volume-title":"Mul-timodal machine learning: A survey and taxonomy","author":"Baltru\u0161aitis T.","year":"2018","unstructured":"T. Baltru\u0161aitis, C. Ahuja, and L.-P. Morency. Mul-timodal machine learning: A survey and taxonomy. IEEE transactions on pattern analysis and machine intelligence, 41(2):423-443, 2018."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNN.2005.853337"}],"container-title":["ACM SIGKDD Explorations Newsletter"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3820356.3820360","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T16:00:41Z","timestamp":1781712041000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3820356.3820360"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,17]]},"references-count":9,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,6,17]]}},"alternative-id":["10.1145\/3820356.3820360"],"URL":"https:\/\/doi.org\/10.1145\/3820356.3820360","relation":{},"ISSN":["1931-0145","1931-0153"],"issn-type":[{"value":"1931-0145","type":"print"},{"value":"1931-0153","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,17]]},"assertion":[{"value":"2026-06-17","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}