{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T06:22:24Z","timestamp":1776838944335,"version":"3.51.2"},"reference-count":57,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2023,5,15]],"date-time":"2023-05-15T00:00:00Z","timestamp":1684108800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Artif. Intell."],"abstract":"<jats:p>Human communication often combines imagery and text into integrated presentations, especially online. In this paper, we show how image\u2013text coherence relations can be used to model the pragmatics of image\u2013text presentations in AI systems. In contrast to alternative frameworks that characterize image\u2013text presentations in terms of the priority, relevance, or overlap of information across modalities, coherence theory postulates that each unit of a discourse stands in specific pragmatic relations to other parts of the discourse, with each relation involving its own information goals and inferential connections. Text accompanying an image may, for example, characterize what's visible in the image, explain how the image was obtained, offer the author's appraisal of or reaction to the depicted situation, and so forth. The advantage of coherence theory is that it provides a simple, robust, and effective abstraction of communicative goals for practical applications. To argue this, we review case studies describing coherence in image\u2013text data sets, predicting coherence from few-shot annotations, and coherence models of image\u2013text tasks such as caption generation and caption evaluation.<\/jats:p>","DOI":"10.3389\/frai.2023.1048874","type":"journal-article","created":{"date-parts":[[2023,5,15]],"date-time":"2023-05-15T05:07:02Z","timestamp":1684127222000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Image\u2013text coherence and its implications for multimodal AI"],"prefix":"10.3389","volume":"6","author":[{"given":"Malihe","family":"Alikhani","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Baber","family":"Khalid","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Matthew","family":"Stone","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2023,5,15]]},"reference":[{"key":"B1","first-page":"9","article-title":"\u201cApplying discourse semantics and pragmatics to co-reference in picture sequences,\u201d","volume-title":"Proceedings of Sinn und Bedeutung 17","author":"Abusch","year":"2013"},{"key":"B2","first-page":"570","article-title":"\u201cCite: a corpus of image-text discourse relations,\u201d","author":"Alikhani","year":"2019","journal-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)"},{"key":"B3","doi-asserted-by":"crossref","first-page":"10427","DOI":"10.1609\/aaai.v36i10.21285","article-title":"\u201cCross-modal coherence for text-to-image retrieval,\u201d","author":"Alikhani","year":"2022","journal-title":"Proceedings of the 36th AAAI Conference on Artificial Intelligence, Vol. 10"},{"key":"B4","doi-asserted-by":"crossref","first-page":"6525","DOI":"10.18653\/v1\/2020.acl-main.583","article-title":"\u201cCross-modal coherence modeling for caption generation,\u201d","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Alikhani","year":"2020"},{"key":"B5","doi-asserted-by":"crossref","first-page":"272","DOI":"10.1109\/MIPR.2018.00063","article-title":"\u201cExploring coherence in visual explanations,\u201d","volume-title":"2018 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR)","author":"Alikhani","year":"2018"},{"key":"B6","doi-asserted-by":"crossref","first-page":"58","DOI":"10.18653\/v1\/W19-1806","article-title":"\u201c\u201ccaption\u201d as a coherence relation: evidence and implications,\u201d","volume-title":"Proceedings of the Second Workshop on Shortcomings in Vision and Language","author":"Alikhani","year":"2019"},{"key":"B7","volume-title":"Logics of Conversation","author":"Asher","year":"2003"},{"key":"B8","doi-asserted-by":"publisher","first-page":"163","DOI":"10.1177\/0261927X00019002001","article-title":"Visible acts of meaning: an integrated message model of language in face-to-face dialogue","volume":"19","author":"Bavelas","year":"2000","journal-title":"J. Lang. Soc. Psychol"},{"key":"B9","doi-asserted-by":"publisher","first-page":"633","DOI":"10.2307\/412039","article-title":"Accent is predictable (if you're a mind-reader)","volume":"48","author":"Bolinger","year":"1972","journal-title":"Language"},{"key":"B10","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1909.00749","article-title":"Know2look: commonsense knowledge for visual search","author":"Chowdhury","year":"2019","journal-title":"arXiv preprint"},{"key":"B11","doi-asserted-by":"publisher","first-page":"413","DOI":"10.1111\/cogs.12016","article-title":"Visual narrative structure","volume":"37","author":"Cohn","year":"2013","journal-title":"Cogn. Sci"},{"key":"B12","first-page":"1","article-title":"Conventions of viewpoint coherence in film","volume":"17","author":"Cumming","year":"2017","journal-title":"Philos. Imprint"},{"key":"B13","volume-title":"Toward a Theory of Multimodal Communication: Combining Speech, Gestures, Diagrams and Demonstrations in Instructional Explanations","author":"Engle","year":"2001"},{"key":"B14","doi-asserted-by":"publisher","first-page":"33","DOI":"10.1109\/2.97249","article-title":"Automating the generation of coordinated multimedia explanations","volume":"24","author":"Feiner","year":"1991","journal-title":"IEEE Comput"},{"key":"B15","doi-asserted-by":"crossref","first-page":"585","DOI":"10.18653\/v1\/D15-1070","article-title":"\u201cImage-mediated learning for zero-shot cross-lingual document retrieval,\u201d","author":"Funaki","year":"2015","journal-title":"Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing"},{"key":"B16","first-page":"4204","article-title":"\u201cMscap: multi-style image captioning with unpaired stylized text,\u201d","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Guo","year":"2019"},{"key":"B17","doi-asserted-by":"publisher","first-page":"661","DOI":"10.1007\/s10579-020-09517-1","article-title":"Ai2d-rst: amultimodal corpus of 1000 primary school science diagrams","volume":"55","author":"Hiippala","year":"2021","journal-title":"Lang. Resour. Evaluat"},{"key":"B18","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1207\/s15516709cog0301_4","article-title":"Coherence and coreference","volume":"3","author":"Hobbs","year":"1979","journal-title":"Cogn. Sci"},{"key":"B19","volume-title":"On the coherence and structure of discourse","author":"Hobbs","year":"1985"},{"key":"B20","first-page":"1233","article-title":"\u201cVisual storytelling,\u201d","author":"Huang","year":"2016","journal-title":"Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies"},{"key":"B21","doi-asserted-by":"publisher","first-page":"317","DOI":"10.1007\/s10988-018-9244-0","article-title":"Relating gesture to speech: reflections on the role of conditional presuppositions","volume":"42","author":"Hunter","year":"2019","journal-title":"Linguist. Philos"},{"key":"B22","first-page":"371","article-title":"\u201cRhetorical relations and quds.\u201d","volume-title":"New Frontiers in Artificial Intelligence: JSAI-isAI Workshops LENLS, JURISIN, KCSD, LLLL Revised Selected Papers","author":"Hunter","year":"2017"},{"key":"B23","doi-asserted-by":"publisher","DOI":"10.3765\/sp.11.10","article-title":"A formal semantics for situated conversation","author":"Hunter","year":"2018","journal-title":"Semant Pragmat"},{"key":"B24","first-page":"3419","article-title":"\u201cCOSMic: a coherence-aware generation metric for image descriptions,\u201d","author":"Inan","year":"2021","journal-title":"Findings of the Association for Computational Linguistics: EMNLP 2021"},{"key":"B25","volume-title":"Coherence, Reference, and the Theory of Grammar","author":"Kehler","year":"2002"},{"key":"B26","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511807572","volume-title":"Gesture: Visible Action as Utterance","author":"Kendon","year":"2004"},{"key":"B27","doi-asserted-by":"crossref","first-page":"4622","DOI":"10.18653\/v1\/D19-1469","article-title":"\u201cIntegrating text and image: Determining multimodal document intent in Instagram posts,\u201d","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)","author":"Kruk","year":"2019"},{"key":"B28","doi-asserted-by":"publisher","first-page":"1956","DOI":"10.1007\/s11263-020-01316-z","article-title":"The open images dataset v4","volume":"128","author":"Kuznetsova","year":"2020","journal-title":"Int. J. Comput. Vis"},{"key":"B29","doi-asserted-by":"publisher","first-page":"147","DOI":"10.1075\/gest.9.2.01las","article-title":"Discourse coherence and gesture interpretation","volume":"9","author":"Lascarides","year":"","journal-title":"Gesture"},{"key":"B30","doi-asserted-by":"publisher","first-page":"393","DOI":"10.1093\/jos\/ffp004","article-title":"A formal semantic analysis of gesture","volume":"26","author":"Lascarides","year":"","journal-title":"J. Semant"},{"key":"B31","first-page":"740","article-title":"\u201cMicrosoft coco: common objects in context,\u201d","volume-title":"European Conference on Computer Vision","author":"Lin","year":"2014"},{"key":"B32","first-page":"13","article-title":"\u201cVilbert: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,\u201d","author":"Lu","year":"2019","journal-title":"Proceedings of the 33rd International Conference on Neural Information Processing Systems"},{"key":"B33","volume-title":"Understanding Comics: The Invisible Art","author":"McCloud","year":"1993"},{"key":"B34","volume-title":"Hand and Mind: What Gestures Reveal About Thought","author":"McNeill","year":"1992"},{"key":"B35","first-page":"15","article-title":"Temporal ontology and temporal reference","volume":"14","author":"Moens","year":"1988","journal-title":"Comput. Linguist"},{"key":"B36","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.conll-1.14","article-title":"Understanding guided image captioning performance across domains","author":"Ng","year":"2020","journal-title":"arXiv preprint"},{"key":"B37","doi-asserted-by":"crossref","first-page":"168","DOI":"10.1145\/3323873.3325049","article-title":"\u201cUnderstanding, categorizing and predicting semantic image-text relations,\u201d","volume-title":"Proceedings of the 2019 on International Conference on Multimedia Retrieval","author":"Otto","year":"2019"},{"key":"B38","first-page":"627","article-title":"\u201cA calculus of cohesion,\u201d","volume-title":"Fourth LACUS Forum","author":"Phillips","year":"1977"},{"key":"B39","article-title":"\u201cThe penn discourse TreeBank 2.0,\u201d","volume-title":"Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC'08)","author":"Prasad","year":"2008"},{"key":"B40","first-page":"8748","article-title":"\u201cLearning transferable visual models from natural language supervision,\u201d","volume-title":"Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research","author":"Radford","year":"2021"},{"key":"B41","first-page":"8821","article-title":"\u201cZero-shot text-to-image generation,\u201d","volume-title":"Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research","author":"Ramesh","year":"2021"},{"key":"B42","doi-asserted-by":"publisher","first-page":"6","DOI":"10.3765\/sp.5.6","article-title":"Information structure: towards an integrated formal theory of pragmatics","volume":"5","author":"Roberts","year":"2012","journal-title":"Semant Pragmat"},{"key":"B43","doi-asserted-by":"crossref","first-page":"2257","DOI":"10.18653\/v1\/P18-1210","article-title":"\u201cDiscourse coherence: Concurrent explicit and implicit relations,\u201d","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Rohde","year":"2018"},{"key":"B44","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1080\/01638539209544800","article-title":"Toward a taxonomy of coherence relations","volume":"15","author":"Sanders","year":"1992","journal-title":"Discour. Process"},{"key":"B45","doi-asserted-by":"crossref","first-page":"40","DOI":"10.18653\/v1\/2020.ecnlp-1.6","article-title":"\u201cImproving intent classification in an E-commerce voice assistant by using inter-utterance context,\u201d","volume-title":"Proceedings of The 3rd Workshop on e-Commerce and NLP","author":"Sharma","year":"2020"},{"key":"B46","doi-asserted-by":"crossref","first-page":"2556","DOI":"10.18653\/v1\/P18-1238","article-title":"\u201cConceptual captions: a cleaned, hypernymed, image alt-text dataset for automatic image captioning,\u201d","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Sharma","year":"2018"},{"key":"B47","first-page":"12516","article-title":"\u201cEngaging image captioning via personality,\u201d","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019","author":"Shuster","year":"2019"},{"key":"B48","doi-asserted-by":"publisher","first-page":"197","DOI":"10.1007\/BF02139085","article-title":"Grammatical determinants of ambiguous pronoun resolution","volume":"23","author":"Smyth","year":"1994","journal-title":"J. Psycholinguist. Res"},{"key":"B49","doi-asserted-by":"publisher","first-page":"69","DOI":"10.1007\/s13164-014-0213-4","article-title":"Meaning and demonstration","volume":"6","author":"Stone","year":"2015","journal-title":"Rev. Philos. Psychol"},{"key":"B50","doi-asserted-by":"publisher","first-page":"109","DOI":"10.1017\/S002222670000058X","article-title":"Discourse structure, topicality and questioning","volume":"31","author":"van Kuppevelt","year":"1995","journal-title":"J. Linguist"},{"key":"B51","first-page":"5998","article-title":"\u201cAttention is all you need,\u201d","author":"Vaswani","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B52","doi-asserted-by":"crossref","first-page":"2830","DOI":"10.18653\/v1\/P19-1272","article-title":"\u201cCategorizing and inferring the relationship between the text and image of Twitter posts,\u201d","volume-title":"Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics","author":"Vempala","year":"2019"},{"key":"B53","author":"Webber","year":"2019","journal-title":"The penn discourse treebank 3.0 annotation manual"},{"key":"B54","doi-asserted-by":"publisher","first-page":"545","DOI":"10.1162\/089120103322753347","article-title":"Anaphora and discourse structure","volume":"29","author":"Webber","year":"2003","journal-title":"Comput. Linguist"},{"key":"B55","doi-asserted-by":"crossref","first-page":"147","DOI":"10.3115\/981175.981196","article-title":"\u201cThe interpretation of tense in discourse,\u201d","volume-title":"Proceedings of the 25th Annual Meeting on Association for Computational Linguistics","author":"Webber","year":"1987"},{"key":"B56","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1166","article-title":"Recipeqa: a challenge dataset for multimodal comprehension of cooking recipes","author":"Yagcioglu","year":"2018","journal-title":"arXiv preprint"},{"key":"B57","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1162\/tacl_a_00166","article-title":"From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions","volume":"2","author":"Young","year":"2014","journal-title":"TACL"}],"container-title":["Frontiers in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2023.1048874\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,15]],"date-time":"2023-05-15T05:07:42Z","timestamp":1684127262000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2023.1048874\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,15]]},"references-count":57,"alternative-id":["10.3389\/frai.2023.1048874"],"URL":"https:\/\/doi.org\/10.3389\/frai.2023.1048874","relation":{},"ISSN":["2624-8212"],"issn-type":[{"value":"2624-8212","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,15]]},"article-number":"1048874"}}