{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,3]],"date-time":"2026-04-03T01:51:39Z","timestamp":1775181099377,"version":"3.50.1"},"publisher-location":"Cham","reference-count":25,"publisher":"Springer Nature Switzerland","isbn-type":[{"value":"9783031712906","type":"print"},{"value":"9783031712913","type":"electronic"}],"license":[{"start":{"date-parts":[[2024,1,1]],"date-time":"2024-01-01T00:00:00Z","timestamp":1704067200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,9,9]],"date-time":"2024-09-09T00:00:00Z","timestamp":1725840000000},"content-version":"vor","delay-in-days":252,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Human-generated captions for photographs, particularly snapshots, have been extensively collected in recent AI research. They play a crucial role in the development of systems capable of multimodal information processing that combines vision and language. Recognizing that diagrams may serve a distinct function in thinking and communication compared to photographs, we shifted our focus from snapshot photographs to diagrams. We provided humans with text-free diagrams and collected data on the captions they generated. The diagrams were sourced from AI2D-RST, a subset of AI2D. This subset annotates the AI2D image dataset of diagrams from elementary school science textbooks with types of diagrams. We mosaicked all textual elements within the diagram images to ensure that human annotators focused solely on the diagram\u2019s visual content when writing a sentence about what the image expresses. For the 831 images in our dataset, we obtained caption data from at least three individuals per image. To the best of our knowledge, this dataset is the first collection of caption data specifically for diagrams.<\/jats:p>","DOI":"10.1007\/978-3-031-71291-3_32","type":"book-chapter","created":{"date-parts":[[2024,9,8]],"date-time":"2024-09-08T14:02:00Z","timestamp":1725804120000},"page":"393-401","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Building a\u00a0Large Dataset of\u00a0Human-Generated Captions for\u00a0Science Diagrams"],"prefix":"10.1007","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4095-7451","authenticated-orcid":false,"given":"Yuri","family":"Sato","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-8978-0325","authenticated-orcid":false,"given":"Ayaka","family":"Suzuki","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2801-9171","authenticated-orcid":false,"given":"Koji","family":"Mineshima","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,9,9]]},"reference":[{"key":"32_CR1","unstructured":"Alikhani M, Stone M: Arrows are the verbs of diagrams. In: COLING 2018, pp. 3552\u20133563. ACL (2018)"},{"key":"32_CR2","unstructured":"Berkeley G.: A Treatise Concerning the Principles of Human Knowledge. The Floating Press (1710\/2014)"},{"key":"32_CR3","doi-asserted-by":"publisher","first-page":"409","DOI":"10.1613\/jair.4900","volume":"55","author":"R Bernardi","year":"2016","unstructured":"Bernardi, R., et al.: Automatic description generation from images: a survey of models, datasets, and evaluation measures. J. Artif. Intell. Res. 55, 409\u2013442 (2016). https:\/\/doi.org\/10.1613\/jair.4900","journal-title":"J. Artif. Intell. Res."},{"key":"32_CR4","doi-asserted-by":"publisher","first-page":"155","DOI":"10.1016\/S0376-6357(01)00156-5","volume":"54","author":"LA Best","year":"2001","unstructured":"Best, L.A., Smith, L.D., Stubbs, D.A.: Graph use in psychology and other sciences. Behav. Process. 54, 155\u2013165 (2001). https:\/\/doi.org\/10.1016\/S0376-6357(01)00156-5","journal-title":"Behav. Process."},{"key":"32_CR5","unstructured":"Daston L, Galison P: Objectivity. Zone Books (2007)"},{"key":"32_CR6","unstructured":"Greenberg, G: Tagging: semantics at the iconic\/symbolic interface. In: AC 2019, pp. 11\u201320. University of Amsterdam (2019)"},{"key":"32_CR7","series-title":"Lecture Notes in Computer Science (Lecture Notes in Artificial Intelligence)","doi-asserted-by":"publisher","first-page":"303","DOI":"10.1007\/978-3-642-31223-6_33","volume-title":"Diagrammatic Representation and Inference","author":"LP Fanjoy","year":"2012","unstructured":"Fanjoy, L.P., MacNeill, A.L., Best, L.A.: The use of diagrams in science. In: Cox, P., Plimmer, B., Rodgers, P. (eds.) Diagrams 2012. LNCS (LNAI), vol. 7352, pp. 303\u2013305. Springer, Heidelberg (2012). https:\/\/doi.org\/10.1007\/978-3-642-31223-6_33"},{"key":"32_CR8","doi-asserted-by":"publisher","first-page":"377","DOI":"10.2307\/2182440","volume":"66","author":"P Grice","year":"1957","unstructured":"Grice, P.: Meaning. Philos. Rev. 66, 377\u2013388 (1957). https:\/\/doi.org\/10.2307\/2182440","journal-title":"Philos. Rev."},{"key":"32_CR9","doi-asserted-by":"publisher","first-page":"661","DOI":"10.1007\/s10579-020-09517-1","volume":"55","author":"T Hiippala","year":"2021","unstructured":"Hiippala, T., et al.: AI2D-RST: a multimodal corpus of 1000 primary school science diagrams. Lang. Resour. Eval. 55, 661\u2013688 (2021). https:\/\/doi.org\/10.1007\/s10579-020-09517-1","journal-title":"Lang. Resour. Eval."},{"key":"32_CR10","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"235","DOI":"10.1007\/978-3-319-46493-0_15","volume-title":"Computer Vision \u2013 ECCV 2016","author":"A Kembhavi","year":"2016","unstructured":"Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., Farhadi, A.: A diagram is worth a dozen images. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV 2016. LNCS, vol. 9908, pp. 235\u2013251. Springer, Cham (2016). https:\/\/doi.org\/10.1007\/978-3-319-46493-0_15"},{"key":"32_CR11","doi-asserted-by":"publisher","first-page":"213","DOI":"10.1002\/acp.2350050305","volume":"5","author":"JA Krosnick","year":"1991","unstructured":"Krosnick, J.A.: Response strategies for coping with the cognitive demands of attitude measures in surveys. Appl. Cogn. Psychol. 5, 213\u2013236 (1991). https:\/\/doi.org\/10.1002\/acp.2350050305","journal-title":"Appl. Cogn. Psychol."},{"key":"32_CR12","unstructured":"Leibniz G.: Philosophical Papers and Letters; Dialogue. L.E. Loemker (Trans. & Ed.). University of Chicago Press, Chicago (1677\/1956)"},{"key":"32_CR13","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"740","DOI":"10.1007\/978-3-319-10602-1_48","volume-title":"Computer Vision \u2013 ECCV 2014","author":"T-Y Lin","year":"2014","unstructured":"Lin, T.-Y., et al.: Microsoft COCO: common objects in context. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014. LNCS, vol. 8693, pp. 740\u2013755. Springer, Cham (2014). https:\/\/doi.org\/10.1007\/978-3-319-10602-1_48"},{"key":"32_CR14","doi-asserted-by":"publisher","first-page":"240","DOI":"10.1037\/0022-0663.81.2.240","volume":"81","author":"RE Mayer","year":"1989","unstructured":"Mayer, R.E.: Systematic thinking fostered by illustrations in scientific text. J. Educ. Psychol. 81, 240\u2013246 (1989). https:\/\/doi.org\/10.1037\/0022-0663.81.2.240","journal-title":"J. Educ. Psychol."},{"key":"32_CR15","doi-asserted-by":"publisher","unstructured":"Mogull S A, Stanfield C T: Current use of visuals in scientific communication. In: IPCC 2015, pp. 1\u20136. IEEE (2015). https:\/\/doi.org\/10.1109\/IPCC.2015.7235818","DOI":"10.1109\/IPCC.2015.7235818"},{"key":"32_CR16","unstructured":"Pauwels, L. (ed.): Visual Cultures of Science: Rethinking Representational Practices in Knowledge Building and Science Communication. Dartmouth College Press (2006)"},{"key":"32_CR17","doi-asserted-by":"publisher","first-page":"103","DOI":"10.1007\/s10339-018-0877-2","volume":"20","author":"Y Sato","year":"2019","unstructured":"Sato, Y., Stapleton, G., Jamnik, M., Shams, Z.: Human inference beyond syllogisms: an approach using external graphical representations. Cogn. Process. 20, 103\u2013115 (2019). https:\/\/doi.org\/10.1007\/s10339-018-0877-2","journal-title":"Cogn. Process."},{"key":"32_CR18","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"373","DOI":"10.1007\/978-3-031-15146-0_34","volume-title":"Diagrammatic Representation and Inference","author":"Y Sato","year":"2022","unstructured":"Sato, Y., Mineshima, K.: Visually analyzing universal quantifiers in photograph captions. In: Giardino, V., Linker, S., Burns, R., Bellucci, F., Boucheix, J.M., Viana, P. (eds.) Diagrams 2022. LNCS, vol. 13462, pp. 373\u2013377. Springer, Cham (2022). https:\/\/doi.org\/10.1007\/978-3-031-15146-0_34"},{"issue":"3","key":"32_CR19","doi-asserted-by":"publisher","first-page":"e13258","DOI":"10.1111\/cogs.13258","volume":"47","author":"Y Sato","year":"2023","unstructured":"Sato, Y., Mineshima, K., Ueda, K.: Can negation be depicted? Comparing human and machine understanding of visual representations. Cogn. Sci. 47(3), e13258 (2023). https:\/\/doi.org\/10.1111\/cogs.13258","journal-title":"Cogn. Sci."},{"key":"32_CR20","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"39","DOI":"10.1007\/978-3-031-55245-8_3","volume-title":"Human and Artificial Rationalities","author":"Y Sato","year":"2024","unstructured":"Sato, Y., Mineshima, K.: Can machines and humans use negation when describing images? In: Baratgin, J., Jacquet, B., Yama, H. (eds.) Human and Artificial Rationalities. Lecture Notes in Computer Science, vol. 14522, pp. 39\u201347. Springer, Cham (2024). https:\/\/doi.org\/10.1007\/978-3-031-55245-8_3"},{"key":"32_CR21","unstructured":"Sato, Y., Suzuki, A., Mineshima, K.: Capturing stage-level and individual-level information from photographs: Human-AI comparison. In: CogSci 2024, pp. 803\u2013810. Cognitive Science Society (2024)"},{"key":"32_CR22","series-title":"Lecture Notes in Computer Science (Lecture Notes in Artificial Intelligence)","doi-asserted-by":"publisher","first-page":"247","DOI":"10.1007\/978-3-319-91376-6_25","volume-title":"Diagrammatic Representation and Inference","author":"Z Shams","year":"2018","unstructured":"Shams, Z., Sato, Y., Jamnik, M., Stapleton, G.: Accessible reasoning with diagrams: from\u00a0cognition to automation. In: Chapman, P., Stapleton, G., Moktefi, A., Perez-Kriz, S., Bellucci, F. (eds.) Diagrams 2018. LNCS (LNAI), vol. 10871, pp. 247\u2013263. Springer, Cham (2018). https:\/\/doi.org\/10.1007\/978-3-319-91376-6_25"},{"key":"32_CR23","doi-asserted-by":"publisher","first-page":"246","DOI":"10.1086\/448287","volume":"11","author":"KL Walton","year":"1984","unstructured":"Walton, K.L.: Transparent pictures. Crit. Inq. 11, 246\u2013277 (1984). https:\/\/doi.org\/10.1086\/448287","journal-title":"Crit. Inq."},{"key":"32_CR24","doi-asserted-by":"crossref","unstructured":"Yoshikawa, Y., Shigeto, Y., Takeuchi, A.: STAIR captions: constructing a large-scale Japanese image caption dataset. In: ACL 2017, pp. 417\u2013421 (2017)","DOI":"10.18653\/v1\/P17-2066"},{"key":"32_CR25","unstructured":"Zala, A., Lin, H., Cho, J., Bansal, M.: DiagrammerGPT: generating open-domain, open-platform diagrams via LLM planning. arXiv:2310.12128 (2023)"}],"container-title":["Lecture Notes in Computer Science","Diagrammatic Representation and Inference"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-031-71291-3_32","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,8]],"date-time":"2024-11-08T13:06:35Z","timestamp":1731071195000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-031-71291-3_32"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"ISBN":["9783031712906","9783031712913"],"references-count":25,"URL":"https:\/\/doi.org\/10.1007\/978-3-031-71291-3_32","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"value":"0302-9743","type":"print"},{"value":"1611-3349","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024]]},"assertion":[{"value":"9 September 2024","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"Diagrams","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"International Conference on Theory and Application of Diagrams","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"M\u00fcnster","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Germany","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2024","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"27 September 2024","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"1 October 2024","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"14","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"diagrams2024","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/diagrams-2024.diagrams-conference.org\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}}]}}