{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T02:24:15Z","timestamp":1780453455370,"version":"3.54.1"},"reference-count":32,"publisher":"Association for Computing Machinery (ACM)","issue":"9","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>\n            Despite interest in multimodal classification, few studies have addressed the missing modality problem in which an\n            <jats:italic toggle=\"yes\">incomplete<\/jats:italic>\n            multimodal input with one or more missing modalities is classified as the target class. The missing modality problem is shown to reduce the classification accuracy as the discriminative power of the obtained feature space is reduced. In this study, we address the missing modality problem in multimodal classification using a novel cascaded framework. The proposed framework is formulated in the feature space to address the missing modality problem by generating\n            <jats:italic toggle=\"yes\">complete<\/jats:italic>\n            multimodal data from\n            <jats:italic toggle=\"yes\">incomplete<\/jats:italic>\n            multimodal data. Subsequently, an optimal multimodal data is obtained by feature selection of the generated and original data. The proposed cascaded framework consists of three steps: feature extraction, feature generation, and classification. The framework is formulated to handle both\n            <jats:italic toggle=\"yes\">complete<\/jats:italic>\n            and\n            <jats:italic toggle=\"yes\">incomplete<\/jats:italic>\n            multimodal data simultaneously. The cascaded framework is trained using novel latent loss functions: missing modality joint loss, centroid joint loss, and latent prior loss. These loss functions, based on metric learning, are designed to ensure that data from the same class remain proximate in the latent space irrespective of the presence or absence of modality data. The cascaded framework is validated on bimodal audio-visible RAVDESS and trimodal audio-visible-thermal Speaking Faces datasets. The experimental results show that the cascaded framework improves classification accuracy even with\n            <jats:italic toggle=\"yes\">incomplete<\/jats:italic>\n            multimodal data.\n          <\/jats:p>","DOI":"10.1145\/3711860","type":"journal-article","created":{"date-parts":[[2025,1,9]],"date-time":"2025-01-09T07:32:25Z","timestamp":1736407945000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Multimodal Cascaded Framework with Multimodal Latent Loss Functions Robust to Missing Modalities"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9553-0906","authenticated-orcid":false,"given":"Vijay","family":"John","sequence":"first","affiliation":[{"name":"Guardian Robot Project, RIKEN, Seika, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3799-4550","authenticated-orcid":false,"given":"Yasutomo","family":"Kawanishi","sequence":"additional","affiliation":[{"name":"Guardian Robot Project, RIKEN, Seika, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,11]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"crossref","unstructured":"Madina Abdrakhmanova Askat Kuzdeuov Sheikh Jarju Yerbolat Khassanov Michael Lewis and Huseyin Atakan Varol. 2020. SpeakingFaces: A large-scale multimodal dataset of voice commands with visual and thermal video streams. arXiv:2012.02961. Retrieved from https:\/\/arxiv.org\/abs\/2012.02961","DOI":"10.3390\/s21103465"},{"key":"e_1_3_1_3_2","first-page":"1","article-title":"A survey on deep multimodal learning for computer vision: Advances, trends, applications, and datasets","volume":"10","author":"Bayoudh Khaled","year":"2021","unstructured":"Khaled Bayoudh, Raja Knani, Fay\u00e7al Hamdaoui, and Abdellatif Mtibaa. 2021. A survey on deep multimodal learning for computer vision: Advances, trends, applications, and datasets. Visual Computing 10 (June 2021), 1\u201332.","journal-title":"Visual Computing"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2006.01.017"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219963"},{"key":"e_1_3_1_6_2","first-page":"17","volume-title":"Proceedings of the 2005 NICTA-HCSNet Multimodal User Interaction Workshop","volume":"57","author":"Chetty Girija","year":"2006","unstructured":"Girija Chetty and Michael Wagner. 2006. Audio-visual multimodal fusion for biometric person authentication and liveness verification. In Proceedings of the 2005 NICTA-HCSNet Multimodal User Interaction Workshop, Vol. 57, 17\u201324."},{"key":"e_1_3_1_7_2","first-page":"176","volume-title":"Proceedings of the International Conference on Audio- and Video-Based Biometric Person Authentication","author":"Choudhury Tanzeem","year":"1999","unstructured":"Tanzeem Choudhury, Brian Clarkson, Tony Jebara, and Alex Pentland. 1999. Multimodal person recognition using unconstrained audio and video. In Proceedings of the International Conference on Audio- and Video-Based Biometric Person Authentication, 176\u2013181."},{"key":"e_1_3_1_8_2","first-page":"605","volume-title":"Proceedings of APSIPA, Annual Summit and Conference","author":"Das Rohan Kumar","year":"2020","unstructured":"Rohan Kumar Das, Ruijie Tao, Jichen Yang, Wei Rao, Cheng Yu, and Haizhou Li. 2020. HLT-NUS submission for 2019 NIST multimedia speaker recognition evaluation. In Proceedings of APSIPA, Annual Summit and Conference, 605\u2013609."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.5555\/3305381.3305478"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8682773"},{"key":"e_1_3_1_11_2","unstructured":"Andrew G. Howard Menglong Zhu Bo Chen Dmitry Kalenichenko Weijun Wang Tobias Weyand MarcoAndreetto and Hartwig Adam. 2017. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv:1704.04861. Retrieved from https:\/\/arxiv.org\/abs\/1704.04861"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3551626.3564965"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3587819.3590989"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-020-08628-9"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-006-6655-0"},{"key":"e_1_3_1_16_2","first-page":"150","volume-title":"Proceedings of the 35th International Conference on Neural Information Processing","author":"Li Qinbo","unstructured":"Qinbo Li, Qing Wan, Sang-Heon Lee, and Yoonsuck Choe. 2021. Video face recognition with audio-visual aggregation network. In Proceedings of the 35th International Conference on Neural Information Processing, 150\u2013161."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0196391"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01764"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/DICTA47822.2019.8945863"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.59543\/ijmscs.v1i.7737"},{"key":"e_1_3_1_21_2","first-page":"400","volume-title":"Proceedings of the 2020 International Conference on Multimodal Interaction","author":"Parthasarathy Srinivas","year":"2020","unstructured":"Srinivas Parthasarathy and Shiva Sundaram. 2020. Training strategies to handle missing modalities for audio-visual expression recognition. In Proceedings of the 2020 International Conference on Multimodal Interaction, 400\u2013404."},{"key":"e_1_3_1_22_2","first-page":"6892","volume-title":"Proceedings of the 33rd AAAI Conference on Artificial Intelligence","author":"Pham Hai","unstructured":"Hai Pham, Paul Pu Liang, Thomas Manzini, Louis-Philippe Morency, and Barnab\u00e1s P\u00f3czos. 2019. Found in translation: Learning robust joint representations by cyclic translations between modalities. In Proceedings of the 33rd AAAI Conference on Artificial Intelligence, 6892\u20136899."},{"key":"e_1_3_1_23_2","unstructured":"Nicolae-Catalin Ristea Liviu-Cristian Dutu and Anamaria Radoi. 2020. Emotion recognition system from speech and visual information based on convolutional neural networks. arXiv:2003.00351. Retrieved from https:\/\/arxiv.org\/abs\/2003.00351"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11036-020-01530-6"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.21437\/Odyssey.2020-37"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1142\/S0218001417560055"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8462122"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1117\/12.543549"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2020-1814"},{"key":"e_1_3_1_30_2","first-page":"1","article-title":"Audio-visual person recognition using deep convolutional neural networks","volume":"8","author":"Vegad Sagar","year":"2017","unstructured":"Sagar Vegad, Harshita Pathak Rajendra Patel, Hanqi Zhuang, and Mehul R. Naik. 2017. Audio-visual person recognition using deep convolutional neural networks. Journal of Biometrics & Biostatistics 8 (2017), 1\u20137.","journal-title":"Journal of Biometrics & Biostatistics"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3366423.3380000"},{"key":"e_1_3_1_32_2","first-page":"1","volume-title":"Proceedings of the 2019 International Conference on Learning Representations","author":"Wen Yandong","year":"2019","unstructured":"Yandong Wen, Mahmoud Al Ismail, Weiyang Liu, Bhiksha Raj, and Rita Singh. 2019. Disjoint mapping network for cross-modal matching of voices and faces. In Proceedings of the 2019 International Conference on Learning Representations, 1\u201317."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.203"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3711860","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,12]],"date-time":"2025-09-12T04:42:18Z","timestamp":1757652138000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3711860"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,11]]},"references-count":32,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3711860"],"URL":"https:\/\/doi.org\/10.1145\/3711860","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,9,11]]},"assertion":[{"value":"2023-12-07","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-10","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}