{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T03:44:05Z","timestamp":1782359045392,"version":"3.54.5"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"11","funder":[{"name":"Science and Technology Innovation (STI) 2030\u2014Major Projects","award":["2022ZD0208700"],"award-info":[{"award-number":["2022ZD0208700"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62376264 and 61702502"],"award-info":[{"award-number":["62376264 and 61702502"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["CUC230B020 and CUC24QT20"],"award-info":[{"award-number":["CUC230B020 and CUC24QT20"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>\n                    Recent multimodal fake news detection methods often use the consistency between textual and visual contents to determine the truth or fake of news information. Higher levels of textual-visual consistency typically lead to a greater likelihood of classifying a news item as real. However, a critical observation reveals that creators of most fake news intentionally select images that align with the textual content, thereby enhancing the credibility of the news. Consequently, high consistency between textual and visual contents alone cannot guarantee the authenticity of the information. To address this problem, we introduce a novel approach termed\n                    <jats:italic toggle=\"yes\">Multimodal Consistency-based Suppression Factor<\/jats:italic>\n                    to modulate the significance of textual-visual consistency in information assessment. When the textual-visual matching is high, this suppression factor reduces the influence of consistency during the judgment process. Moreover, we use contrastive language-image pre-training (CLIP) model to extract features and measure the consistency level between modalities to guide multimodal fusion. In addition, we also use a method of compressing and fusing modal information based on variational autoencoder (VAE) to reconstruct CLIP features, learning the shared representation of different modal information of CLIP. Finally, extensive experiments were conducted on three publicly datasets, Weibo, Twitter, and Weibo21, and the results confirmed that our method outperformed the state-of-the-art methods in the field and had 0.8%, 2.6%, and 4.1% effect improvement on the accuracy rate.\n                  <\/jats:p>","DOI":"10.1145\/3699959","type":"journal-article","created":{"date-parts":[[2024,10,14]],"date-time":"2024-10-14T08:27:19Z","timestamp":1728894439000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Multimodal Consistency Suppression Factor for Fake News Detection"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9011-8464","authenticated-orcid":false,"given":"Zhulin","family":"Tao","sequence":"first","affiliation":[{"name":"School of Information and Communication and Engineering, Communication University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-2161-8408","authenticated-orcid":false,"given":"Runze","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Computer and Cyber Sciences, Communication University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7699-9514","authenticated-orcid":false,"given":"Xin","family":"Shi","sequence":"additional","affiliation":[{"name":"School of Information and Communication and Engineering, Communication University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4660-8092","authenticated-orcid":false,"given":"Xingyu","family":"Gao","sequence":"additional","affiliation":[{"name":"Institute of Microelectronics, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6642-8160","authenticated-orcid":false,"given":"Xi","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Microelectronics, Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0324-4687","authenticated-orcid":false,"given":"Xianglin","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Computer and Cyber Sciences, Communication University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,11,10]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1257\/jep.31.2.211"},{"key":"e_1_3_1_3_2","first-page":"999","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Bhunia Ayan Kumar","year":"2022","unstructured":"Ayan Kumar Bhunia, Subhadeep Koley, Abdullah Faiz Ur Rahman Khilji, Aneeshan Sain, Pinaki Nath Chowdhury, Tao Xiang, and Yi-Zhe Song. 2022. Sketching without worrying: Noise-tolerant sketch-based image retrieval. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 999\u20131008."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i01.5393"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13735-017-0143-x"},{"key":"e_1_3_1_6_2","first-page":"1877","article-title":"Language models are few-shot learners","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems, 1877\u20131901.","journal-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems"},{"key":"e_1_3_1_7_2","first-page":"6848","volume-title":"Proceedings of the Language Resources and Evaluation Conference","author":"Carlsson Fredrik","year":"2022","unstructured":"Fredrik Carlsson, Philipp Eisen, Faton Rekathati, and Magnus Sahlgren. 2022. Cross-lingual and multilingual CLIP. In Proceedings of the Language Resources and Evaluation Conference. European Language Resources Association, 6848\u20136854."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01392"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3485447.3511968"},{"key":"e_1_3_1_10_2","first-page":"104","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920)","author":"Chen Yen-Chun","year":"2020","unstructured":"Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020. Uniter: Universal image-text representation learning. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u201920). Springer, 104\u2013120."},{"key":"e_1_3_1_11_2","unstructured":"Aakanksha Chowdhery Sharan Narang Jacob Devlin Maarten Bosma Gaurav Mishra Adam Roberts Paul Barham Hyung Won Chung Charles Sutton Sebastian Gehrmann et al. 2022. Palm: Scaling language modeling with pathways. arXiv:2204.02311. Retrieved from https:\/\/arxiv.org\/abs\/2204.02311"},{"key":"e_1_3_1_12_2","first-page":"3951","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)","author":"Conde Marcos V.","year":"2021","unstructured":"Marcos V. Conde and Kerem Turgutlu. 2021. CLIP-Art: Contrastive pre-training for fine-grained art classification. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 3951\u20133955."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1002\/pra2.2015.145052010082"},{"key":"e_1_3_1_14_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Retrieved from https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i1.16080"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3463001"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01028"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123454"},{"key":"e_1_3_1_19_2","first-page":"1546","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics","author":"Karimi Hamid","year":"2018","unstructured":"Hamid Karimi, Proteek Roy, Sari Saba-Sadiya, and Jiliang Tang. 2018. Multi-source multi-class fake news detection. In Proceedings of the 27th International Conference on Computational Linguistics, 1546\u20131557."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313552"},{"key":"e_1_3_1_21_2","first-page":"5583","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Kim Wonjae","year":"2021","unstructured":"Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In Proceedings of the International Conference on Machine Learning. PMLR, 5583\u20135594."},{"key":"e_1_3_1_22_2","unstructured":"Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv:1312.6114. Retrieved from https:\/\/arxiv.org\/abs\/1312.6114"},{"key":"e_1_3_1_23_2","first-page":"1103","volume-title":"Proceedings of the IEEE 13th International Conference on Data Mining","author":"Kwon Sejeong","year":"2013","unstructured":"Sejeong Kwon, Meeyoung Cha, Kyomin Jung, Wei Chen, and Yajun Wang. 2013. Prominent features of rumor propagation in online social media. In Proceedings of the IEEE 13th International Conference on Data Mining. IEEE, 1103\u20131108."},{"key":"e_1_3_1_24_2","unstructured":"Junnan Li Dongxu Li Silvio Savarese and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv:2301.12597. Retrieved from https:\/\/arxiv.org\/abs\/2301.12597"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3617827"},{"key":"e_1_3_1_26_2","unstructured":"Han Liu Yinwei Wei Xuemeng Song Weili Guan Yuan-Fang Li and Liqiang Nie. 2024. MMGRec: Multimodal generative recommendation with transformer model. arxiv:2404.16555. Retrieved from https:\/\/arxiv.org\/abs\/2404.16555"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/2806416.2806651"},{"key":"e_1_3_1_28_2","first-page":"3818","article-title":"Detecting rumors from microblogs with recurrent neural networks","author":"Ma Jing","year":"2016","unstructured":"Jing Ma, Wei Gao, Prasenjit Mitra, Sejeong Kwon, Bernard J Jansen, Kam-Fai Wong, and Meeyoung Cha. 2016. Detecting rumors from microblogs with recurrent neural networks. In Proceedings of the 25th International Joint Conference on Artificial Intelligence, 3818\u20133824.","journal-title":"Proceedings of the 25th International Joint Conference on Artificial Intelligence"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/2806416.2806607"},{"issue":"3","key":"e_1_3_1_30_2","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1111\/hir.12311","article-title":"The Covid-19 \u2018infodemic\u2019: A new front for information professionals","volume":"37","author":"Naeem Salman Bin","year":"2020","unstructured":"Salman Bin Naeem and Rubina Bhatti. 2020. The Covid-19 \u2018infodemic\u2019: A new front for information professionals. Health Information & Libraries Journal 37, 3 (2020), 233\u2013239.","journal-title":"Health Information & Libraries Journal"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3459637.3482139"},{"key":"e_1_3_1_32_2","first-page":"518","article-title":"Exploiting multi-domain visual information for fake news detection","author":"Qi Peng","year":"2019","unstructured":"Peng Qi, Juan Cao, Tianyun Yang, Junbo Guo, and Jintao Li. 2019. Exploiting multi-domain visual information for fake news detection. In Proceedings of the IEEE International Conference on Data Mining (ICDM). IEEE, 518\u2013527.","journal-title":"Proceedings of the IEEE International Conference on Data Mining (ICDM)"},{"key":"e_1_3_1_33_2","first-page":"8748","article-title":"Learning transferable visual models from natural language supervision","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, 8748\u20138763.","journal-title":"Proceedings of the International Conference on Machine Learning"},{"key":"e_1_3_1_34_2","unstructured":"Aditya Ramesh Prafulla Dhariwal Alex Nichol Casey Chu and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv:2204.06125. Retrieved from https:\/\/arxiv.org\/abs\/2204.06125"},{"key":"e_1_3_1_35_2","first-page":"28","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems 28 (2015).","journal-title":"Proceedings of the 28th International Conference on Neural Information Processing Systems"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3487553.3524650"},{"key":"e_1_3_1_38_2","first-page":"39","article-title":"Spotfake: A multi-modal framework for fake news detection","author":"Singhal Shivangi","year":"2019","unstructured":"Shivangi Singhal, Rajiv Ratn Shah, Tanmoy Chakraborty, Ponnurangam Kumaraguru, and Shin\u2019ichi Satoh. 2019. Spotfake: A multi-modal framework for fake news detection. In Proceedings of the IEEE 5th International Conference on Multimedia Big Data (BigMM). IEEE, 39\u201347.","journal-title":"Proceedings of the IEEE 5th International Conference on Multimedia Big Data (BigMM)"},{"key":"e_1_3_1_39_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et al. 2023. Llama: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_1_40_2","first-page":"11","article-title":"Visualizing data using t-SNE","volume":"9","author":"Maaten Laurens Van der","year":"2008","unstructured":"Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research 9, 11 (2008), 2579\u20132605.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20059-5_40"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219903"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01754"},{"key":"e_1_3_1_44_2","unstructured":"Chuhan Wu Fangzhao Wu Tao Qi Yongfeng Huang and Xing Xie. 2022. NoisyTune: A little noise can help you finetune pretrained language models better. arXiv:2202.12024. Retrieved from https:\/\/arxiv.org\/abs\/2202.12024"},{"key":"e_1_3_1_45_2","first-page":"2560","article-title":"Multimodal fusion with co-attention networks for fake news detection","author":"Wu Yang","year":"2021","unstructured":"Yang Wu, Pengwei Zhan, Yunjian Zhang, Liming Wang, and Zhen Xu. 2021. Multimodal fusion with co-attention networks for fake news detection. In Proceedings of the Findings of the Association for Computational Linguistics (ACL-IJCNLP \u201921), 2560\u20132569.","journal-title":"Proceedings of the Findings of the Association for Computational Linguistics (ACL-IJCNLP \u201921)"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2021.102610"},{"key":"e_1_3_1_47_2","first-page":"3901","article-title":"A convolutional approach for misinformation identification","author":"Yu Feng","year":"2017","unstructured":"Feng Yu, Qiang Liu, Shu Wu, Liang Wang, and Tieniu Tan. 2017. A convolutional approach for misinformation identification. In Proceedings of the 26th International Joint Conference on Artificial Intelligence (IJCAI), 3901\u20133907.","journal-title":"Proceedings of the 26th International Joint Conference on Artificial Intelligence (IJCAI)"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442381.3450004"},{"key":"e_1_3_1_49_2","first-page":"2413","volume-title":"Proceedings of the 31st International Joint Conference on Artificial Intelligence (IJCAI \u201922)","author":"Zheng Jiaqi","year":"2022","unstructured":"Jiaqi Zheng, Xi Zhang, Sanchuan Guo, Quan Wang, Wenyu Zang, and Yongdong Zhang. 2022. MFAN: Multi-modal feature-enhanced attention networks for rumor detection. In Proceedings of the 31st International Joint Conference on Artificial Intelligence (IJCAI \u201922), 2413\u20132419."},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1007\/978-3-030-47436-2_27","volume-title":"Proceedings of the 24th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining (PAKDD \u201920)","author":"Zhou Xinyi","year":"2020","unstructured":"Xinyi Zhou, Jindi Wu, and Reza Zafarani. 2020. Similarity-aware multi-modal fake news detection. In Proceedings of the 24th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining (PAKDD \u201920), 354\u2013367."},{"key":"e_1_3_1_51_2","unstructured":"Deyao Zhu Jun Chen Xiaoqian Shen Xiang Li and Mohamed Elhoseiny. 2023. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. arXiv:2304.10592. Retrieved from https:\/\/arxiv.org\/abs\/2304.10592"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3161603"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3699959","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,10]],"date-time":"2025-11-10T14:51:06Z","timestamp":1762786266000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3699959"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,10]]},"references-count":51,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3699959"],"URL":"https:\/\/doi.org\/10.1145\/3699959","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,10]]},"assertion":[{"value":"2024-04-26","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-09-23","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}