{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,13]],"date-time":"2026-07-13T20:31:52Z","timestamp":1783974712695,"version":"3.55.0"},"reference-count":151,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2023,5,2]],"date-time":"2023-05-02T00:00:00Z","timestamp":1682985600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Artif. Intell."],"abstract":"<jats:p>The analysis of news dissemination is of utmost importance since the credibility of information and the identification of disinformation and misinformation affect society as a whole. Given the large amounts of news data published daily on the Web, the empirical analysis of news with regard to research questions and the detection of problematic news content on the Web require computational methods that work at scale. Today's online news are typically disseminated in a multimodal form, including various presentation modalities such as text, image, audio, and video. Recent developments in multimodal machine learning now make it possible to capture basic \u201cdescriptive\u201d relations between modalities\u2013such as correspondences between words and phrases, on the one hand, and corresponding visual depictions of the verbally expressed information on the other. Although such advances have enabled tremendous progress in tasks like image captioning, text-to-image generation and visual question answering, in domains such as news dissemination, there is a need to go further. In this paper, we introduce a novel framework for the computational analysis of multimodal news. We motivate a set of more complex image-text relations as well as multimodal news values based on real examples of news reports and consider their realization by computational approaches. To this end, we provide (a) an overview of existing literature from <jats:italic>semiotics<\/jats:italic> where detailed proposals have been made for taxonomies covering diverse image-text relations generalisable to any domain; (b) an overview of computational work that derives models of image-text relations from data; and (c) an overview of a particular class of news-centric attributes developed in journalism studies called news values. The result is a novel framework for multimodal news analysis that closes existing gaps in previous work while maintaining and combining the strengths of those accounts. We assess and discuss the elements of the framework with real-world examples and use cases, setting out research directions at the intersection of multimodal learning, multimodal analytics and computational social sciences that can benefit from our approach.<\/jats:p>","DOI":"10.3389\/frai.2023.1125533","type":"journal-article","created":{"date-parts":[[2023,5,2]],"date-time":"2023-05-02T04:54:18Z","timestamp":1683003258000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":17,"title":["Understanding image-text relations and news values for multimodal news analysis"],"prefix":"10.3389","volume":"6","author":[{"given":"Gullal S.","family":"Cheema","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sherzod","family":"Hakimov","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eric","family":"M\u00fcller-Budack","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christian","family":"Otto","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"John A.","family":"Bateman","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ralph","family":"Ewerth","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2023,5,2]]},"reference":[{"key":"B1","first-page":"1","article-title":"\u201cAnalyzing user modeling on twitter for personalized news recommendations,\u201d","volume-title":"User Modeling, Adaption and Personalization - 19th International Conference, UMAP 2011","author":"Abel","year":"2011"},{"key":"B2","first-page":"2962","article-title":"\u201cTwitter-based user modeling for news recommendations,\u201d","volume-title":"IJCAI 2013, Proceedings of the 23rd International Joint Conference on Artificial Intelligence","author":"Abel","year":"2013"},{"key":"B3","first-page":"6139","article-title":"\u201cFact vs. opinion: the role of argumentation features in news classification,\u201d","volume-title":"Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020","author":"Alhindi","year":"2020"},{"key":"B4","doi-asserted-by":"publisher","first-page":"6525","DOI":"10.18653\/v1\/2020.acl-main.583","article-title":"\u201cCross-modal coherence modeling for caption generation,\u201d","author":"Alikhani","year":"2020","journal-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics"},{"key":"B5","author":"Aneja","year":"2021"},{"key":"B6","doi-asserted-by":"publisher","first-page":"633","DOI":"10.1177\/1464884918809299","article-title":"News values on social media: Exploring what drives peaks in user activity about organizations on twitter","volume":"21","author":"Araujo","year":"2020","journal-title":"Journalism"},{"key":"B7","doi-asserted-by":"crossref","first-page":"3154","DOI":"10.18653\/v1\/2020.acl-main.287","article-title":"\u201cAnalyzing the persuasive effect of style in news editorial argumentation,\u201d","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020","author":"Baff","year":"2020"},{"key":"B8","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1109\/TPAMI.2018.2798607","article-title":"Multimodal machine learning: A survey and taxonomy","volume":"41","author":"Baltrusaitis","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"B9","volume-title":"Image-Music-Text","author":"Barthes","year":"1977"},{"key":"B10","doi-asserted-by":"crossref","DOI":"10.4324\/9781315773971","volume-title":"Text and Image: A Critical Introduction to the Visual\/Verbal Divide","author":"Bateman","year":"2014"},{"key":"B11","doi-asserted-by":"publisher","first-page":"227","DOI":"10.3366\/cor.2016.0093","article-title":"Investigating evaluation and news values in news items that are shared through social media","volume":"11","author":"Bednarek","year":"2016","journal-title":"Corpora"},{"key":"B12","doi-asserted-by":"publisher","first-page":"103","DOI":"10.1016\/j.dcm.2012.05.006","article-title":"\u201cvalue added\u201d: Language, image and news values","volume":"1","author":"Bednarek","year":"2012","journal-title":"Discour. Context Media"},{"key":"B13","doi-asserted-by":"crossref","DOI":"10.1093\/acprof:oso\/9780190653934.001.0001","volume-title":"The Discourse of News Values: How News Organizations Create Newsworthiness","author":"Bednarek","year":"2017"},{"key":"B14","doi-asserted-by":"publisher","first-page":"702","DOI":"10.1080\/1461670X.2020.1807393","article-title":"Computer-based analysis of news values: A case study on national day reporting","volume":"22","author":"Bednarek","year":"2021","journal-title":"Journal. Stud"},{"key":"B15","volume-title":"The Language of News Media","author":"Bell","year":"1991"},{"key":"B16","doi-asserted-by":"publisher","first-page":"1132","DOI":"10.31449\/inf.v42i4.1132","article-title":"Automatic estimation of news values reflecting importance and closeness of news events","volume":"42","author":"Belyaeva","year":"2018","journal-title":"Informatica"},{"key":"B17","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511621024","volume-title":"Variation Across Speech and Writing","author":"Biber","year":"1988"},{"key":"B18","doi-asserted-by":"crossref","DOI":"10.4135\/9781446216026","volume-title":"News Values","author":"Brighton","year":"2007"},{"key":"B19","first-page":"5410","article-title":"Image-text retrieval: A survey on recent research and development,\u201d","volume-title":"Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI 2022","author":"Cao","year":"2022"},{"key":"B20","doi-asserted-by":"crossref","DOI":"10.1057\/9781137314901","volume-title":"Photojournalism: A Social Semiotic Approach","author":"Caple","year":"2013"},{"key":"B21","doi-asserted-by":"publisher","first-page":"435","DOI":"10.1177\/1464884914568078","article-title":"Rethinking news values: What a discursive approach can tell us about the construction of news discourse and news photography","volume":"17","author":"Caple","year":"2016","journal-title":"Journalism"},{"key":"B22","volume-title":"DNVA and Intratextual Analysis","author":"Caple","year":"2017"},{"key":"B23","doi-asserted-by":"publisher","DOI":"10.1017\/9781108886048","author":"Caple","year":"2020","journal-title":"Multimodal News Analysis across Cultures"},{"key":"B24","doi-asserted-by":"crossref","first-page":"77","DOI":"10.18653\/v1\/W17-2711","article-title":"\u201cThe event storyline corpus: A new benchmark for causal and temporal relation extraction,\u201d","volume-title":"Proceedings of the Events and Stories in the News Workshop@ACL 2017","author":"Caselli","year":"2017"},{"key":"B25","first-page":"781","article-title":"\u201cUnderstanding and classifying image tweets,\u201d","volume-title":"ACM Multimedia Conference, MM '13","author":"Chen","year":"2013"},{"key":"B26","doi-asserted-by":"crossref","first-page":"104","DOI":"10.1007\/978-3-030-58577-8_7","article-title":"\u201cUNITER: universal image-text representation learning,\u201d","volume-title":"Computer Vision - ECCV 2020 - 16th European Conference","author":"Chen","year":"2020"},{"key":"B27","doi-asserted-by":"publisher","first-page":"10","DOI":"10.1186\/s40537-022-00561-y","article-title":"Part of speech tagging: a systematic review of deep learning and machine learning approaches","volume":"9","author":"Chiche","year":"2022","journal-title":"J. Big Data"},{"key":"B28","doi-asserted-by":"crossref","first-page":"663","DOI":"10.18653\/v1\/D19-1061","article-title":"\u201cExtracting possessions from social media: Images complement language,\u201d","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019","author":"Chinnappa","year":"2019"},{"key":"B29","doi-asserted-by":"crossref","first-page":"2833","DOI":"10.18653\/v1\/2021.findings-emnlp.242","article-title":"\u201cBe nice to your wife! the restaurants are closed\u201d: Can gender stereotype detection improve sexism classification?,\u201d","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event","author":"Chiril","year":"2021"},{"key":"B30","doi-asserted-by":"publisher","first-page":"273","DOI":"10.1007\/BF00994018","article-title":"Support-vector networks","volume":"20","author":"Cortes","year":"1995","journal-title":"Mach. Learn"},{"key":"B31","first-page":"248","article-title":"\u201cImagenet: A large-scale hierarchical image database,\u201d","volume-title":"2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009)","author":"Deng","year":"2009"},{"key":"B32","first-page":"4171","article-title":"\u201cBERT: pre-training of deep bidirectional transformers for language understanding,\u201d","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019","author":"Devlin","year":"2019"},{"key":"B33","doi-asserted-by":"crossref","first-page":"1","DOI":"10.18653\/v1\/W17-4201","article-title":"\u201cPredicting news values from headline text and emotions,\u201d","volume-title":"Proceedings of the 2017 Workshop: Natural Language Processing meets Journalism, NLPmJ@EMNLP","author":"di Buono","year":"2017"},{"key":"B34","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3479550","article-title":"Towards understanding and supporting journalistic practices using semi-automated news discovery tools","volume":"5","author":"Diakopoulos","year":"2021","journal-title":"Proc. ACM Human-Comput. Inter"},{"key":"B35","author":"D'Ignazio","year":"2014","journal-title":"Cliff-clavin: Determining geographic focus for news articles"},{"key":"B36","article-title":"\u201cStudying muslim stereotyping through microportrait extraction,\u201d","volume-title":"Proceedings of the Eleventh International Conference on Language Resources and Evaluation, LREC 2018","author":"Fokkens","year":"2018"},{"key":"B37","doi-asserted-by":"publisher","first-page":"64","DOI":"10.1177\/002234336500200104","article-title":"The structure of foreign news: The presentation of the congo, cuba and cyprus crises in four norwegian newspapers","volume":"2","author":"Galtung","year":"1965","journal-title":"J. Peace Res"},{"key":"B38","doi-asserted-by":"publisher","first-page":"163","DOI":"10.1561\/0600000105","article-title":"Vision-language pre-training: Basics, recent advances, and future trends","volume":"14","author":"Gan","year":"2022","journal-title":"Found. Trends Comput. Graph. Vis"},{"key":"B39","first-page":"30","article-title":"\u201cMultimodal fake news detection with textual, visual and semantic information,\u201d","volume-title":"Text, Speech, and Dialogue - 23rd International Conference, TSD 2020","author":"Giachanou","year":"2020"},{"key":"B40","article-title":"\u201cLarge-scale sentiment analysis for news and blogs,\u201d","volume-title":"Proceedings of the First International Conference on Weblogs and Social Media, ICWSM 2007","author":"Godbole","year":"2007"},{"key":"B41","first-page":"17","article-title":"Fake news vs satire: A dataset and analysis,\u201d","volume-title":"Proceedings of the 10th ACM Conference on Web Science, WebSci 2018","author":"Golbeck","year":"2018"},{"key":"B42","article-title":"Bertopic: Neural topic modeling with a class-based TF-IDF procedure. CoRR, abs\/2203.05794","author":"Grootendorst","year":"2022"},{"key":"B43","first-page":"6047","article-title":"AVA: A video dataset of spatio-temporally localized atomic visual actions,\u201d","volume-title":"2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018","author":"Gu","year":"2018"},{"key":"B44","doi-asserted-by":"publisher","first-page":"22","DOI":"10.1016\/j.neucom.2020.02.139","article-title":"Deep learning-based aerial image segmentation with open data for disaster impact assessment","volume":"439","author":"Gupta","year":"2021","journal-title":"Neurocomputing"},{"key":"B45","volume-title":"An Introduction to Functional Grammar","author":"Halliday","year":"1985"},{"key":"B46","doi-asserted-by":"crossref","DOI":"10.4324\/9780203783771","volume-title":"An Introduction to Functional Grammar","author":"Halliday","year":"2014"},{"key":"B47","first-page":"1859","article-title":"\u201cA retrospective analysis of the fake news challenge stance-detection task,\u201d","volume-title":"Proceedings of the 27th International Conference on Computational Linguistics, COLING 2018","author":"Hanselowski","year":"2018"},{"key":"B48","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1080\/14616700118449","article-title":"What is news? Galtung and ruge revisited","volume":"2","author":"Harcup","year":"2001","journal-title":"Journal. Stud"},{"key":"B49","doi-asserted-by":"publisher","first-page":"1470","DOI":"10.1080\/1461670X.2016.1150193","article-title":"What is news? News values revisited (again)","volume":"18","author":"Harcup","year":"2017","journal-title":"Journal. Stud"},{"key":"B50","first-page":"961","article-title":"\u201cActivitynet: A large-scale video benchmark for human activity understanding,\u201d","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015","author":"Heilbron","year":"2015"},{"key":"B51","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1007\/s13735-017-0142-y","article-title":"Estimating the information gap between textual and visual representations","volume":"7","author":"Henning","year":"2018","journal-title":"Int. J. Multim. Inf. Retr"},{"key":"B52","doi-asserted-by":"publisher","first-page":"377","DOI":"10.1177\/0270467610385893","article-title":"The presentation of self in the age of social media: Distinguishing performances and exhibitions online","volume":"30","author":"Hogan","year":"2010","journal-title":"Bull. Sci. Technol. Soc"},{"key":"B53","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3295748","article-title":"A comprehensive survey of deep learning for image captioning","volume":"51","author":"Hossain","year":"2019","journal-title":"ACM Comput. Surv"},{"key":"B54","doi-asserted-by":"crossref","first-page":"1956","DOI":"10.1109\/BigData.2017.8258141","article-title":"\u201cFocus location extraction from political news reports with bias correction,\u201d","volume-title":"2017 IEEE International Conference on Big Data (IEEE BigData 2017)","author":"Imani","year":"2017"},{"key":"B55","first-page":"4904","article-title":"\u201cScaling up visual and vision-language representation learning with noisy text supervision,\u201d","volume-title":"Proceedings of the 38th International Conference on Machine Learning, ICML 2021","author":"Jia","year":"2021"},{"key":"B56","doi-asserted-by":"publisher","first-page":"157","DOI":"10.17645\/mac.v7i3.1910","article-title":"Newsworthiness and the public's response in russian social media: A comparison of state and private news organizations","volume":"7","author":"Judina","year":"2019","journal-title":"Media Communic"},{"key":"B57","doi-asserted-by":"publisher","first-page":"177","DOI":"10.1080\/21670811.2015.1096619","article-title":"Content analysis and online news: epistemologies of analysing the ephemeral web","volume":"4","author":"Karlsson","year":"2016","journal-title":"Digital Journal"},{"key":"B58","first-page":"3128","article-title":"\u201cDeep visual-semantic alignments for generating image descriptions,\u201d","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015","author":"Karpathy","year":"2015"},{"key":"B59","doi-asserted-by":"publisher","first-page":"18167","DOI":"10.1007\/s11042-019-08571-4","article-title":"Estimating the imageability of words by mining visual characteristics from crawled image data","volume":"79","author":"Kastner","year":"2020","journal-title":"Multim. Tools Appl"},{"key":"B60","first-page":"1351","article-title":"\u201cPatterns of argumentation strategies across topics,\u201d","volume-title":"Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017","author":"Khatib","year":"2017"},{"key":"B61","article-title":"Unifying visual-semantic embeddings with multimodal neural language models. CoRR, abs\/1411.2539","author":"Kiros","year":"2014"},{"key":"B62","first-page":"42","article-title":"Komplementarit\u00e4t von sprache und bild am beispiel von comic, karikatur und reklame.(la compl\u00e9mentarit\u00e9 de la langue et de l'image. l'exemple des bandes dessin\u00e9es, des caricatures et des r\u00e9clames)","volume":"57","author":"Kloepfer","year":"1976","journal-title":"Sprache Techn. Zeitalter Stuttgart"},{"key":"B63","doi-asserted-by":"publisher","first-page":"687","DOI":"10.1017\/S1351324917000043","article-title":"Classifying news versus opinions in newspapers: Linguistic features for domain independence","volume":"23","author":"Kr\u00fcger","year":"2017","journal-title":"Nat. Lang. Eng"},{"key":"B64","doi-asserted-by":"crossref","first-page":"4621","DOI":"10.18653\/v1\/D19-1469","article-title":"\u201cIntegrating text and image: Determining multimodal document intent in instagram posts,\u201d","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019","author":"Kruk","year":"2019"},{"key":"B65","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"B66","first-page":"87","article-title":"Multiplying meaning: visual and verbal semiotics in scientific text,\u201d","volume-title":"Reading science: critical and functional perspectives on discourses of science","author":"Lemke","year":"1998"},{"key":"B67","doi-asserted-by":"publisher","first-page":"1195","DOI":"10.1109\/TAFFC.2020.2981446","article-title":"Deep facial expression recognition: A survey","volume":"13","author":"Li","year":"2022","journal-title":"IEEE Trans. Affect. Comput"},{"key":"B68","doi-asserted-by":"publisher","first-page":"367","DOI":"10.1109\/TMM.2016.2616279","article-title":"Joint image-text news topic detection and tracking by multimodal topic and-or graph","volume":"19","author":"Li","year":"2017","journal-title":"IEEE Trans. Multim"},{"key":"B69","doi-asserted-by":"crossref","first-page":"6761","DOI":"10.18653\/v1\/2021.emnlp-main.542","article-title":"\u201cVisual news: Benchmark and challenges in news image captioning,\u201d","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021","author":"Liu","year":"2021"},{"key":"B70","doi-asserted-by":"crossref","first-page":"6801","DOI":"10.18653\/v1\/2021.emnlp-main.545","article-title":"\u201cNewsclippings: Automatic generation of out-of-context multimodal media,\u201d","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021","author":"Luo","year":"2021"},{"key":"B71","doi-asserted-by":"crossref","first-page":"879","DOI":"10.18653\/v1\/D15-1104","article-title":"\u201cJoint entity recognition and disambiguation,\u201d","volume-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015","author":"Luo","year":"2015"},{"key":"B72","doi-asserted-by":"publisher","first-page":"3339","DOI":"10.1145\/2858036.2858160","article-title":"\u201cConstructing the visual online political self: an analysis of instagram use by the scottish electorate,\u201d","author":"Mahoney","year":"2016","journal-title":"Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems"},{"key":"B73","article-title":"\u201cGenerating images from captions with attention,\u201d","volume-title":"4th International Conference on Learning Representations, ICLR 2016","author":"Mansimov","year":"2016"},{"key":"B74","doi-asserted-by":"publisher","first-page":"647","DOI":"10.1108\/00220410310506303","article-title":"A taxonomy of relationships between images and text","volume":"59","author":"Marsh","year":"2003","journal-title":"J. Document"},{"key":"B75","first-page":"29","article-title":"Macro-genres: the ecology of the page","volume":"21","author":"Martin","year":"1994","journal-title":"Network"},{"key":"B76","volume-title":"Genre Relations: Mapping Culture","author":"Martin","year":"2008"},{"key":"B77","doi-asserted-by":"publisher","first-page":"337","DOI":"10.1177\/1470357205055928","article-title":"A system for image-text relations in new (and old) media","volume":"4","author":"Martinec","year":"2005","journal-title":"Visual Communic"},{"key":"B78","article-title":"\u201cSocial media semantics: Analysing meanings in multimodal online conversations,\u201d","volume-title":"Proceedings of the International Conference on Information Systems - Building a Better World through Information Systems, ICIS 2014","author":"Mehmet","year":"2014"},{"key":"B79","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s42803-022-00052-9","article-title":"Combining sentiment analysis classifiers to explore multilingual news articles covering london 2012 and rio 2016 olympics","volume":"10","author":"Mello","year":"2022","journal-title":"Int. J. Digital Human"},{"key":"B80","doi-asserted-by":"publisher","first-page":"626","DOI":"10.3758\/BF03192732","article-title":"Emotional category data on images from the international affective picture system","volume":"37","author":"Mikels","year":"2005","journal-title":"Behav. Res. Methods"},{"key":"B81","first-page":"23","article-title":"\u201cGenre as social action,\u201d","volume-title":"Genre and the New Rhetoric, Chapter 2","author":"Miller","year":"1994"},{"key":"B82","doi-asserted-by":"publisher","first-page":"120613","DOI":"10.1109\/ACCESS.2020.3005513","article-title":"Analysis and design of computational news angles","volume":"8","author":"Motta","year":"2020","journal-title":"IEEE Access"},{"key":"B83","volume-title":"A Multimodal Analysis of Picture Books for Children: A Systemic Functional Approach","author":"Moya Guijarro","year":"2014"},{"key":"B84","first-page":"619","article-title":"\u201cWhen was this picture taken? Image date estimation in the wild,\u201d","volume-title":"Advances in Information Retrieval - 39th European Conference on IR Research, ECIR 2017","author":"M\u00fcller","year":"2017"},{"key":"B85","doi-asserted-by":"publisher","first-page":"575","DOI":"10.1007\/978-3-030-01258-8_35","article-title":"\u201cGeolocation estimation of photos using a hierarchical model and scene classification,\u201d","author":"M\u00fcller-Budack","year":"2018","journal-title":"Computer Vision - ECCV 2018 - 15th European Conference"},{"key":"B86","first-page":"2927","article-title":"Ontology-driven event type classification in images,\u201d","volume-title":"IEEE Winter Conference on Applications of Computer Vision, WACV 2021","author":"M\u00fcller-Budack","year":""},{"key":"B87","doi-asserted-by":"publisher","first-page":"111","DOI":"10.1007\/s13735-021-00207-4","article-title":"Multimodal news analytics using measures of cross-modal entity and context consistency","volume":"10","author":"M\u00fcller-Budack","year":"","journal-title":"Int. J. Multim. Inf. Retr"},{"key":"B88","first-page":"689","article-title":"\u201cMultimodal deep learning,\u201d","volume-title":"Proceedings of the 28th International Conference on Machine Learning, ICML 2011","author":"Ngiam","year":"2011"},{"key":"B89","doi-asserted-by":"publisher","first-page":"4372","DOI":"10.25073\/2525-2445\/vnufs.4372","article-title":"Exploring text-image relations in english comics for children: The case of \u201clittle red riding hood\u201d","volume":"35","author":"Nhat","year":"2019","journal-title":"VNU J. Foreign Stud"},{"key":"B90","doi-asserted-by":"publisher","first-page":"100467","DOI":"10.1016\/j.dcm.2021.100467","article-title":"Multimodal approach to analysing big social and news media data","volume":"40","author":"O'Halloran","year":"2021","journal-title":"Discourse, Context Media"},{"key":"B91","doi-asserted-by":"publisher","first-page":"1440","DOI":"10.1049\/iet-ipr.2019.1270","article-title":"Survey on visual sentiment analysis","volume":"14","author":"Ortis","year":"2020","journal-title":"IET Image Process"},{"key":"B92","first-page":"711","article-title":"\u201cIs this an example image?\u201d Predicting the relative abstractness level of image and text,","volume-title":"Advances in Information Retrieval - 41st European Conference on IR Research, ECIR 2019","author":"Otto","year":""},{"key":"B93","first-page":"168","article-title":"Understanding, categorizing and predicting semantic image-text relations,\u201d","volume-title":"Proceedings of the 2019 on International Conference on Multimedia Retrieval, ICMR 2019","author":"Otto","year":""},{"key":"B94","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1007\/s13735-019-00187-6","article-title":"Characterization and classification of semantic image-text relations","volume":"9","author":"Otto","year":"2020","journal-title":"Int. J. Multim. Inf. Retr"},{"key":"B95","first-page":"2855","article-title":"\u201cCrisscrossed captions: Extended intramodal and intermodal semantic similarity judgments for MS-COCO,\u201d","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL 2021","author":"Parekh","year":"2021"},{"key":"B96","doi-asserted-by":"publisher","first-page":"14648849211019895","DOI":"10.1177\/14648849211019895","article-title":"Applying news values theory to liking, commenting and sharing mainstream news articles on facebook","volume":"24","author":"Park","year":"2021","journal-title":"Journalism"},{"key":"B97","doi-asserted-by":"publisher","first-page":"64","DOI":"10.18653\/v1\/E17-4007","article-title":"Automatic extraction of news values from headline text,\u201d","author":"Piotrkowicz","year":"2017","journal-title":"Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2017"},{"key":"B98","doi-asserted-by":"publisher","first-page":"647","DOI":"10.1075\/prag.21.4.07pol","article-title":"Detecting contrast patterns in newspaper articles by combining discourse analysis and text mining","volume":"21","author":"Pollak","year":"2011","journal-title":"Pragmatics"},{"key":"B99","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/P17-1081","article-title":"\u201cContext-dependent sentiment analysis in user-generated videos,\u201d","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017","author":"Poria","year":"2017"},{"key":"B100","doi-asserted-by":"publisher","first-page":"149","DOI":"10.1177\/1750481314568548","article-title":"How can computer-based methods help researchers to investigate news values in large datasets? A corpus linguistic study of the construction of newsworthiness in the reporting on hurricane katrina","volume":"9","author":"Potts","year":"2015","journal-title":"Discour. Commun"},{"key":"B101","first-page":"1505","article-title":"\u201cMirrorgan: Learning text-to-image generation by redescription,\u201d","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019","author":"Qiao","year":"2019"},{"key":"B102","first-page":"8748","article-title":"\u201cLearning transferable visual models from natural language supervision,\u201d","volume-title":"Proceedings of the 38th International Conference on Machine Learning, ICML 2021","author":"Radford","year":"2021"},{"key":"B103","first-page":"8821","article-title":"\u201cZero-shot text-to-image generation,\u201d","volume-title":"Proceedings of the 38th International Conference on Machine Learning, ICML 2021","author":"Ramesh","year":"2021"},{"key":"B104","first-page":"5136","article-title":"Multimodal news article analysis,\u201d","volume-title":"Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI 2017","author":"Ramisa","year":"2017"},{"key":"B105","doi-asserted-by":"crossref","first-page":"2050","DOI":"10.1145\/3297280.3297481","article-title":"\u201cA computationally efficient multi-modal classification approach of disaster-related twitter images,\u201d","volume-title":"Proceedings of the 34th ACM\/SIGAPP Symposium on Applied Computing, SAC 2019","author":"Rizk","year":"2019"},{"key":"B106","first-page":"25","article-title":"Synergy on the page: Exploring intersemiotic complementarity in page-based multimodal text","volume":"1","author":"Royce","year":"1998","journal-title":"JASFL Occas"},{"key":"B107","doi-asserted-by":"publisher","first-page":"3610","DOI":"10.3390\/app11083610","article-title":"How do you speak about immigrants? Taxonomy and stereoimmigrants dataset for identifying stereotypes about immigrants","volume":"11","author":"S\u00e1nchez-Junquera","year":"2021","journal-title":"Appl. Sci"},{"key":"B108","doi-asserted-by":"publisher","first-page":"21503","DOI":"10.1007\/s00521-021-06086-4","article-title":"Predicting image credibility in fake news over social media using multi-modal approach","volume":"34","author":"Singh","year":"2021","journal-title":"Neural Comput. Applic"},{"key":"B109","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1002\/asi.24359","article-title":"Detecting fake news stories via multimodal analysis","volume":"72","author":"Singh","year":"2021","journal-title":"J. Assoc. Inf. Sci. Technol"},{"key":"B110","doi-asserted-by":"publisher","first-page":"1349","DOI":"10.1109\/34.895972","article-title":"Content-based image retrieval at the end of the early years","volume":"22","author":"Smeulders","year":"2000","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"B111","doi-asserted-by":"publisher","first-page":"207","DOI":"10.1162\/tacl_a_00177","article-title":"Grounded compositional semantics for finding and describing images with sentences","volume":"2","author":"Socher","year":"2014","journal-title":"Trans. Assoc. Comput. Linguist"},{"key":"B112","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1016\/j.imavis.2017.08.003","article-title":"A survey of multimodal sentiment analysis","volume":"65","author":"Soleymani","year":"2017","journal-title":"Image Vis. Comput"},{"key":"B113","article-title":"\u201cUsing the image-text relationship to improve multimodal disaster tweet classification,\u201d","author":"Sosea","year":"2021","journal-title":"The 18th International Conference on Information Systems for Crisis Response and Management (ISCRAM 2021)"},{"key":"B114","doi-asserted-by":"crossref","first-page":"2575","DOI":"10.1145\/3404835.3462796","article-title":"\u201cQuti! quantifying text-image consistency in multimodal documents,\u201d","volume-title":"SIGIR '21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Springstein","year":"2021"},{"key":"B115","volume-title":"Textstil und Semiotik englischsprachiger Anzeigenwerbung","author":"St\u00f6ckl","year":"1997"},{"key":"B116","doi-asserted-by":"crossref","DOI":"10.4324\/9780429487965","volume-title":"Shifts Towards Image-Centricity in Contemporary Multimodal Practices","author":"St\u00f6ckl","year":"2020"},{"key":"B117","volume-title":"Genre Analysis: English in Academic and Research Settings","author":"Swales","year":"1990"},{"key":"B118","doi-asserted-by":"crossref","first-page":"2565","DOI":"10.1145\/3404835.3462786","article-title":"Geowine: Geolocation based wiki, image, news and event retrieval,\u201d","volume-title":"SIGIR '21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Tahmasebzadeh","year":"2021"},{"key":"B119","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-28238-6_14","article-title":"Mm-locate-news: Multimodal focus location estimation in news","author":"Tahmasebzadeh","year":"2022"},{"key":"B120","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/ICOMET.2019.8673428","author":"Taj","year":"2019"},{"key":"B121","doi-asserted-by":"publisher","first-page":"110","DOI":"10.17645\/mac.v9i1.3331","article-title":"What is (fake) news? Analyzing news values (and more) in fake stories","volume":"9","author":"Tandoc","year":"2021","journal-title":"Media Communic"},{"key":"B122","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/978-3-030-95073-6_14","article-title":"\u201cDeep learning to encourage citizen involvement in local journalism,\u201d","volume-title":"Futures of Journalism: Technology-stimulated Evolution in the Audience-News Media Relationship","author":"Tessem","year":"2022"},{"key":"B123","doi-asserted-by":"crossref","first-page":"1474","DOI":"10.1109\/WACV51458.2022.00154","article-title":"\u201cInterpretable semantic photo geolocation,\u201d","volume-title":"IEEE\/CVF Winter Conference on Applications of Computer Vision, WACV 2022","author":"Theiner","year":"2022"},{"key":"B124","doi-asserted-by":"publisher","first-page":"64","DOI":"10.1145\/2812802","article-title":"YFCC100M: the new data in multimedia research","volume":"59","author":"Thomee","year":"2016","journal-title":"Commun. ACM"},{"key":"B125","doi-asserted-by":"publisher","first-page":"585","DOI":"10.1007\/s43681-021-00126-4","article-title":"Responsible media technology and ai: challenges and research directions","volume":"2","author":"Trattner","year":"2021","journal-title":"AI Ethics"},{"key":"B126","article-title":"Image\/text relations and intersemiosis: Towards multimodal text description for multiliteracies education,\u201d","volume-title":"Proceedings of the 33rd IFSC: International Systemic Functional Congress","author":"Unsworth","year":"2007"},{"key":"B127","first-page":"53","article-title":"What did this castle look like before? exploring referential relations in naturally occurring multimodal texts,\u201d","volume-title":"Proceedings of the Third Workshop on Beyond Vision and LANguage: inTEgrating Real-world kNowledge (LANTERN)","author":"Utescher","year":"2021"},{"key":"B128","doi-asserted-by":"publisher","first-page":"76","DOI":"10.1080\/10304319109388216","article-title":"Conjunctive structure in documentary film and television","volume":"5","author":"van Leeuwen","year":"1991","journal-title":"Continuum J. Media Cult. Stud"},{"key":"B129","volume-title":"Introducing Social Semiotics","author":"van Leeuwen","year":"2005"},{"key":"B130","first-page":"2830","article-title":"\u201cCategorizing and inferring the relationship between the text and image of twitter posts,\u201d","volume-title":"Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019","author":"Vempala","year":"2019"},{"key":"B131","first-page":"2576","article-title":"\u201cNPA: neural news recommendation with personalized attention,\u201d","volume-title":"Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &Data Mining, KDD 2019","author":"Wu","year":"2019"},{"key":"B132","first-page":"1624","article-title":"\u201cUser-as-graph: User modeling with heterogeneous graph pooling for news recommendation,\u201d","volume-title":"Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI 2021","author":"Wu","year":"2021"},{"key":"B133","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3530257","article-title":"Personalized news recommendation: Methods and challenges","volume":"41","author":"Wu","year":"","journal-title":"ACM Trans. Inf. Syst"},{"key":"B134","doi-asserted-by":"publisher","first-page":"3023","DOI":"10.24963\/ijcai.2020\/418","article-title":"\u201cUser modeling with click preference and reading satisfaction for news recommendation,\u201d","author":"Wu","year":"2020","journal-title":"Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI 2020"},{"key":"B135","first-page":"2560","article-title":"\u201cMm-rec: Visiolinguistic model empowered multimodal news recommendation,\u201d","volume-title":"SIGIR '22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Wu","year":""},{"key":"B136","doi-asserted-by":"publisher","first-page":"1415","DOI":"10.4304\/tpls.4.7.1415-1420","article-title":"A multimodal analysis of image-text relations in picture books","volume":"4","author":"Wu","year":"2014","journal-title":"Theory Pract. Langu. Stud"},{"key":"B137","first-page":"59","article-title":"Winfried n\u00f6th, handbook of semiotics","volume":"111","author":"Wunderli","year":"1995","journal-title":"Zeitschrift Romanische Philol"},{"key":"B138","doi-asserted-by":"crossref","first-page":"3485","DOI":"10.1109\/CVPR.2010.5539970","article-title":"\u201cSUN database: Large-scale scene recognition from abbey to zoo,\u201d","volume-title":"The Twenty-Third IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2010","author":"Xiao","year":"2010"},{"key":"B139","first-page":"1600","article-title":"\u201cRecognize complex events from static images by fusing deep channels,\u201d","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015","author":"Xiong","year":"2015"},{"key":"B140","article-title":"Multimodal learning with transformers: A survey. CoRR, abs\/2206.06488","author":"Xu","year":"2022"},{"key":"B141","first-page":"2346","article-title":"\u201cJointly modeling deep video and compositional text to bridge vision and language in a unified framework,\u201d","volume-title":"Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence","author":"Xu","year":"2015"},{"key":"B142","first-page":"427","article-title":"\u201cSemantic correlation mining between images and texts with global semantics and local mapping,\u201d","volume-title":"MultiMedia Modeling - 21st International Conference, MMM 2015","author":"Xue","year":"2015"},{"key":"B143","doi-asserted-by":"crossref","first-page":"419","DOI":"10.1145\/1101149.1101241","article-title":"\u201cImage region entropy: a measure of \u201cvisualness\u201d of web images associated with one concept,\u201d","volume-title":"Proceedings of the 13th ACM International Conference on Multimedia","author":"Yanai","year":"2005"},{"key":"B144","doi-asserted-by":"publisher","first-page":"9225","DOI":"10.1109\/ACCESS.2018.2886366","article-title":"A novel hot topic detection framework with integration of image and short text information from twitter","volume":"7","author":"Zhang","year":"2019","journal-title":"IEEE Access"},{"key":"B145","article-title":"\u201cEqual but not the same: Understanding the implicit relationship between persuasive images and text,\u201d","volume-title":"British Machine Vision Conference 2018, BMVC 2018","author":"Zhang","year":"2018"},{"key":"B146","first-page":"1945","article-title":"\u201cLearning the semantic correlation: An alternative way to gain from unlabeled text,\u201d","volume-title":"Advances in Neural Information Processing Systems 21, Proceedings of the Twenty-Second Annual Conference on Neural Information Processing Systems","author":"Zhang","year":"2008"},{"key":"B147","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2302.05543","article-title":"Adding conditional control to text-to-image diffusion models","author":"Zhang","year":"2023","journal-title":"arXiv [Preprint].arXiv: 2302.05543"},{"key":"B148","first-page":"10394","article-title":"\u201cDeep supervised cross-modal retrieval,\u201d","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019","author":"Zhen","year":"2019"},{"key":"B149","doi-asserted-by":"publisher","first-page":"1452","DOI":"10.1109\/TPAMI.2017.2723009","article-title":"Places: A 10 million image database for scene recognition","volume":"40","author":"Zhou","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"B150","doi-asserted-by":"crossref","first-page":"741","DOI":"10.1145\/2393347.2396301","article-title":"\u201cGeo-location inference on news articles via multimodal plsa,\u201d","volume-title":"Proceedings of the 20th ACM Multimedia Conference, MM'12","author":"Zhou","year":"2012"},{"key":"B151","doi-asserted-by":"publisher","first-page":"10492","DOI":"10.1109\/CVPR46437.2021.01035","article-title":"Webface260m: A benchmark unveiling the power of million-scale deep face recognition,\u201d","author":"Zhu","year":"2021","journal-title":"IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021"}],"container-title":["Frontiers in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2023.1125533\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,2]],"date-time":"2023-05-02T04:55:22Z","timestamp":1683003322000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2023.1125533\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,2]]},"references-count":151,"alternative-id":["10.3389\/frai.2023.1125533"],"URL":"https:\/\/doi.org\/10.3389\/frai.2023.1125533","relation":{},"ISSN":["2624-8212"],"issn-type":[{"value":"2624-8212","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,2]]},"article-number":"1125533"}}