{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,19]],"date-time":"2026-02-19T05:11:44Z","timestamp":1771477904551,"version":"3.50.1"},"reference-count":37,"publisher":"Oxford University Press (OUP)","issue":"2","license":[{"start":{"date-parts":[[2025,9,24]],"date-time":"2025-09-24T00:00:00Z","timestamp":1758672000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/pages\/standard-publication-reuse-rights"}],"funder":[{"DOI":"10.13039\/501100017691","name":"Key Research and Development Program of Guangxi","doi-asserted-by":"publisher","award":["GuikeAB22080065"],"award-info":[{"award-number":["GuikeAB22080065"]}],"id":[{"id":"10.13039\/501100017691","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation","doi-asserted-by":"crossref","award":["72164006"],"award-info":[{"award-number":["72164006"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation","doi-asserted-by":"crossref","award":["61762029"],"award-info":[{"award-number":["61762029"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Innovation Project of Guangxi Graduate Education","award":["YCSW2023328"],"award-info":[{"award-number":["YCSW2023328"]}]},{"name":"Innovation Project of GUET Graduate Education","award":["2024YCXS090"],"award-info":[{"award-number":["2024YCXS090"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026,2,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Multimodal sentiment analysis utilizes different types of information to improve sentiment recognition, driving research at the intersection of natural language processing and computer vision. However, practical applications often encounter missing image modal data due to collection restrictions, privacy protection, and other issues, which reduces the accuracy of sentiment analysis. Recently, diffusion models, as an emerging generative model, have demonstrated remarkable capabilities in handling modality completion and generation tasks. To this end, this paper presents an expert feature-aligned diffusion model (EFADM) integrated with curriculum learning, aiming to efficiently solve multimodal sentiment analysis tasks by dynamically generating missing modality data and optimizing feature alignment between modalities. EFADM consists of three primary modules: the Adaptive Noise Suppression Module (ANSM), the Expert Alignment Module (EAM), and a progressive generation framework based on curriculum learning. First, ANSM dynamically adjusts noise suppression during the diffusion process to enhance the quality of generated features. EAM strengthens semantic consistency between generated and existing modalities by aligning expert features. Additionally, the progressive generation strategy utilizing curriculum learning gradually increases task difficulty, thereby optimizing the model\u2019s performance across different noise levels. Compared with existing methods, the proposed approach achieves leading performance on the publicly available MVSA series datasets.<\/jats:p>","DOI":"10.1093\/comjnl\/bxaf110","type":"journal-article","created":{"date-parts":[[2025,9,24]],"date-time":"2025-09-24T23:12:32Z","timestamp":1758755552000},"page":"221-235","source":"Crossref","is-referenced-by-count":0,"title":["Multimodal sentiment analysis based on expert feature-aligned diffusion under uncertain image modality"],"prefix":"10.1093","volume":"69","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1569-4538","authenticated-orcid":false,"given":"Pingshan","family":"Liu","sequence":"first","affiliation":[{"name":"School of Business , Guilin University of Electronic Technology, Guilin, 541004,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-3459-2150","authenticated-orcid":false,"given":"Fei","family":"Sun","sequence":"additional","affiliation":[{"name":"School of Business , Guilin University of Electronic Technology, Guilin, 541004,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-6395-2388","authenticated-orcid":false,"given":"Liya","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Business , Guilin University of Electronic Technology, Guilin, 541004,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weiping","family":"Fang","sequence":"additional","affiliation":[{"name":"School of Business , Guilin University of Electronic Technology, Guilin, 541004,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0985-1042","authenticated-orcid":false,"given":"Fu","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Business , Guilin University of Electronic Technology, Guilin, 541004,","place":["China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2025,9,24]]},"reference":[{"key":"2026021823461766700_ref1","doi-asserted-by":"publisher","first-page":"3107","DOI":"10.1093\/comjnl\/bxac153","article-title":"Chinese RoBERTa distillation for emotion classification","volume":"66","author":"Liu","year":"2023","journal-title":"Comput J"},{"key":"2026021823461766700_ref2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.4018\/IJSWIS.359985","article-title":"Semantic-driven crossmodal fusion for multimodal sentiment analysis","volume":"20","author":"Liu","year":"2024","journal-title":"Int J Semant Web Inf Syst"},{"key":"2026021823461766700_ref3","doi-asserted-by":"publisher","first-page":"1856","DOI":"10.1109\/TAFFC.2024.3378570","article-title":"Contrastive learning based modality-invariant feature acquisition for robust multimodal emotion recognition with missing modalities","volume":"15","author":"Liu","year":"2024","journal-title":"IEEE Trans Affective Comput"},{"key":"2026021823461766700_ref4","first-page":"2608","article-title":"Missing modality imagination network for emotion recognition with uncertain missing modalities","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL) and the 11th International Joint Conference on Natural Language Processing (IJCNLP)","author":"Zhao"},{"key":"2026021823461766700_ref5","doi-asserted-by":"crossref","first-page":"5780","DOI":"10.1145\/3664647.3680653","article-title":"Robust multimodal sentiment analysis of image-text pairs by distribution-based feature recovery and fusion","volume-title":"Proceedings of the 32nd ACM International Conference on Multimedia (MM)","author":"Wu","year":"2024"},{"key":"2026021823461766700_ref6","first-page":"1","article-title":"Text-guided reconstruction network for sentiment analysis with uncertain missing modalities","volume":"01","author":"Shi","year":"2025","journal-title":"IEEE Trans Affect Comput"},{"key":"2026021823461766700_ref7","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1145\/3689096.3689460","article-title":"Improved generation of synthetic imaging data using feature-aligned diffusion","volume-title":"Proceedings of the First International Workshop on Vision-Language Models for Biomedical Applications (VLM4Bio)","author":"Nair","year":"2024"},{"key":"2026021823461766700_ref8","first-page":"55943","article-title":"Towards robust multimodal sentiment analysis with incomplete data","volume-title":"Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS)","author":"Zhang","year":"2024"},{"key":"2026021823461766700_ref9","first-page":"3588","article-title":"Multimodal multi-loss fusion network for sentiment analysis","volume-title":"Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL)","author":"Wu","year":"2024"},{"key":"2026021823461766700_ref10","doi-asserted-by":"publisher","first-page":"1686","DOI":"10.1162\/tacl_a_00628","article-title":"MissModal: increasing robustness to missing modality in multimodal sentiment analysis","volume":"11","author":"Lin","year":"2023","journal-title":"Trans Assoc Comput Ling"},{"key":"2026021823461766700_ref11","doi-asserted-by":"crossref","first-page":"4400","DOI":"10.1145\/3474085.3475585","article-title":"Transformer-based feature reconstruction network for robust multimodal sentiment analysis","volume-title":"Proceedings of the 29th ACM International Conference on Multimedia (MM)","author":"Yuan","year":"2021"},{"key":"2026021823461766700_ref12","first-page":"367","article-title":"Conformer: local features coupling global representations for visual recognition","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Peng","year":"2021"},{"key":"2026021823461766700_ref13","first-page":"5301","article-title":"CTFN: Hierarchical learning for multimodal sentiment analysis using coupled-translation fusion network","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics(ACL)","author":"Tang","year":"2021"},{"key":"2026021823461766700_ref14","first-page":"1545","article-title":"Tag-assisted multimodal sentiment analysis under uncertain missing modalities","volume-title":"Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR)","author":"Zeng","year":"2022"},{"key":"2026021823461766700_ref15","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1109\/TAFFC.2023.3274829","article-title":"Efficient multimodal transformer with dual-level feature restoration for robust multimodal sentiment analysis","volume":"15","author":"Sun","year":"2023","journal-title":"IEEE Trans Affect Comput"},{"key":"2026021823461766700_ref16","doi-asserted-by":"crossref","first-page":"102454","DOI":"10.1016\/j.inffus.2024.102454","article-title":"Similar modality completion-based multimodal sentiment analysis under uncertain missing modalities","volume":"110","author":"Sun","year":"2024","journal-title":"Inf Fusion"},{"key":"2026021823461766700_ref17","first-page":"2256","article-title":"Deep unsupervised learning using nonequilibrium thermodynamics","volume-title":"Proceedings of the 32nd International Conference on International Conference on Machine Learning (ICML)","author":"Sohl-Dickstein","year":"2015"},{"key":"2026021823461766700_ref18","first-page":"6840","article-title":"Denoising diffusion probabilistic models","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Ho","year":"2020"},{"key":"2026021823461766700_ref19","first-page":"8162","article-title":"Improved denoising diffusion probabilistic models","volume-title":"Proceedings of the 38th International Conference on Machine Learning (ICML)","author":"Nichol","year":"2021"},{"key":"2026021823461766700_ref20","article-title":"Score-based generative modeling through stochastic differential equations","author":"Song","year":"2020"},{"key":"2026021823461766700_ref21","first-page":"4328","article-title":"Diffusion-LM improves controllable text generation","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Li","year":"2022"},{"key":"2026021823461766700_ref22","first-page":"1336","article-title":"Diffusion models for implicit image segmentation ensembles","volume-title":"Proceedings of the International Conference on Medical Imaging with Deep Learning (MIDL)","author":"Wolleb","year":"2022"},{"key":"2026021823461766700_ref23","first-page":"10684","article-title":"High-resolution image synthesis with latent diffusion models","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Rombach","year":"2022"},{"key":"2026021823461766700_ref24","first-page":"5551","article-title":"Effective integration of text diffusion and pre-trained language models with linguistic easy-first schedule","volume-title":"Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC)","author":"Ou","year":"2024"},{"key":"2026021823461766700_ref25","article-title":"Prompt-to-prompt image editing with cross attention control","author":"Hertz","year":"2022"},{"key":"2026021823461766700_ref26","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1145\/1553374.1553380","article-title":"Curriculum learning","volume-title":"Proceedings of the 26th Annual International Conference on Machine Learning (ICML)","author":"Bengio","year":"2009"},{"key":"2026021823461766700_ref27","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1007\/978-3-319-27674-8_2","article-title":"Sentiment analysis on multi-view social data","volume-title":"MultiMedia Modeling: 22nd International Conference, MMM 2016, Proceedings, Part II (MMM)","author":"Niu","year":"2016"},{"key":"2026021823461766700_ref28","doi-asserted-by":"crossref","first-page":"2399","DOI":"10.1145\/3132847.3133142","article-title":"MultiSentiNet: a deep semantic network for multimodal sentiment analysis","volume-title":"Proceedings of the 2017 ACM on Conference on Information and Knowledge Management (CIKM)","author":"Xu","year":"2017"},{"key":"2026021823461766700_ref29","doi-asserted-by":"crossref","first-page":"2924","DOI":"10.18653\/v1\/2022.emnlp-main.189","article-title":"Mitigating inconsistencies in multimodal sentiment analysis under uncertain missing modalities","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Zeng","year":"2022"},{"key":"2026021823461766700_ref30","doi-asserted-by":"crossref","first-page":"411","DOI":"10.1007\/978-3-031-27818-1_34","article-title":"Multimodal reconstruct and align net for missing modality problem in sentiment analysis","volume-title":"International Conference on Multimedia Modeling (MMM)","author":"Luo","year":"2023"},{"key":"2026021823461766700_ref31","first-page":"2282","article-title":"CLMLF: a contrastive learning and multi-layer fusion method for multimodal sentiment detection","volume-title":"Findings of the Association for Computational Linguistics (NAACL)","author":"Li","year":"2022"},{"key":"2026021823461766700_ref32","first-page":"5240","article-title":"Tackling modality heterogeneity with multi-view calibration network for multimodal sentiment detection","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL)","author":"Wei","year":"2023"},{"key":"2026021823461766700_ref33","doi-asserted-by":"publisher","first-page":"101973","DOI":"10.1016\/j.inffus.2023.101973","article-title":"Modality translation-based multimodal sentiment analysis under uncertain missing modalities","volume":"101","author":"Liu","year":"2024","journal-title":"Inf Fusion"},{"key":"2026021823461766700_ref34","first-page":"19730","article-title":"BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models","volume-title":"International Conference on Machine Learning (ICML), PMLR","author":"Li","year":"2023"},{"key":"2026021823461766700_ref35","first-page":"6840","article-title":"Denoising diffusion probabilistic models","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems (NIPS)","author":"Ho","year":"2020"},{"key":"2026021823461766700_ref36","article-title":"Denoising diffusion implicit models","author":"Song","year":"2020"},{"key":"2026021823461766700_ref37","first-page":"1","article-title":"DPM-solver++: fast solver for guided sampling of diffusion probabilistic models","author":"Lu","year":"2025","journal-title":"Mach Intell Res"}],"container-title":["The Computer Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/69\/2\/221\/64377115\/bxaf110.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/comjnl\/article-pdf\/69\/2\/221\/64377115\/bxaf110.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,19]],"date-time":"2026-02-19T04:46:29Z","timestamp":1771476389000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/comjnl\/article\/69\/2\/221\/8263178"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,24]]},"references-count":37,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,9,24]]},"published-print":{"date-parts":[[2026,2,15]]}},"URL":"https:\/\/doi.org\/10.1093\/comjnl\/bxaf110","relation":{},"ISSN":["0010-4620","1460-2067"],"issn-type":[{"value":"0010-4620","type":"print"},{"value":"1460-2067","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026,2]]},"published":{"date-parts":[[2025,9,24]]}}}