{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,20]],"date-time":"2026-02-20T03:04:58Z","timestamp":1771556698694,"version":"3.50.1"},"reference-count":32,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2023,6,6]],"date-time":"2023-06-06T00:00:00Z","timestamp":1686009600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,6,6]],"date-time":"2023-06-06T00:00:00Z","timestamp":1686009600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Discov Artif Intell"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We present a detailed case study evaluating selective cognitive abilities (decision making and spatial reasoning) of two recently released generative transformer models, ChatGPT and DALL-E 2. Input prompts were constructed following neutral a priori guidelines, rather than adversarial intent. Post hoc qualitative analysis of the outputs shows that DALL-E 2 is able to generate at least one correct image for each spatial reasoning prompt, but most images generated are incorrect, even though the model seems to have a clear understanding of the objects mentioned in the prompt. Similarly, in evaluating ChatGPT on the rationality axioms developed under the classical Von Neumann-Morgenstern utility theorem, we find that, although it demonstrates some level of rational decision-making, many of its decisions violate at least one of the axioms even under reasonable constructions of preferences, bets, and decision-making prompts. ChatGPT\u2019s outputs on such problems generally tended to be unpredictable: even as it made irrational decisions (or employed an incorrect reasoning process) for some simpler decision-making problems, it was able to draw correct conclusions for more complex bet structures. We briefly comment on the nuances and challenges involved in scaling up such a \u2018cognitive\u2019 evaluation or conducting it with a closed set of answer keys (\u2018ground truth\u2019), given that these models are inherently generative and open-ended in responding to prompts.<\/jats:p>","DOI":"10.1007\/s44163-023-00067-3","type":"journal-article","created":{"date-parts":[[2023,6,6]],"date-time":"2023-06-06T18:03:21Z","timestamp":1686074601000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Evaluating deep generative models on cognitive tasks: a case study"],"prefix":"10.1007","volume":"3","author":[{"given":"Zhisheng","family":"Tang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mayank","family":"Kejriwal","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,6,6]]},"reference":[{"key":"67_CR1","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN et al. Attention is all you need. Adv Neural Informat Process Syst. 2017; 30."},{"key":"67_CR2","unstructured":"Devlin J, Chang MW, Lee K, Toutanova K. Bert: Pre-training of deep bidirectional transformers for language understanding. 2018; arXiv preprint arXiv:1810.04805."},{"key":"67_CR3","unstructured":"Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D et al. Roberta: A robustly optimized bert pretraining approach. 2019; arXiv preprint arXiv:1907.11692."},{"key":"67_CR4","doi-asserted-by":"crossref","unstructured":"Bianchi F, Kalluri P, Durmus E, Ladhak F, Cheng M, Nozza D et al. Easily accessible text-to-image generation amplifies demographic stereotypes at large scale. 2022; arXiv preprint arXiv:2211.03759.","DOI":"10.1145\/3593013.3594095"},{"key":"67_CR5","doi-asserted-by":"publisher","first-page":"34","DOI":"10.1162\/tacl_a_00298","volume":"8","author":"A Ettinger","year":"2020","unstructured":"Ettinger A. What BERT is not: lessons from a new suite of psycholinguistic diagnostics for language models. Trans Assoc Comput Linguist. 2020;8:34\u201348.","journal-title":"Trans Assoc Comput Linguist"},{"key":"67_CR6","doi-asserted-by":"publisher","first-page":"842","DOI":"10.1162\/tacl_a_00349","volume":"8","author":"A Rogers","year":"2021","unstructured":"Rogers A, Kovaleva O, Rumshisky A. A primer in BERTology: what we know about how BERT works. Trans Assoc Comput Linguist. 2021;8:842\u201366.","journal-title":"Trans Assoc Comput Linguist"},{"key":"67_CR7","doi-asserted-by":"crossref","unstructured":"Gessler L, Schneider N. BERT has uncommon sense: Similarity ranking for word sense BERTology. 2021; arXiv preprint arXiv:2109.09780.","DOI":"10.18653\/v1\/2021.blackboxnlp-1.43"},{"key":"67_CR8","unstructured":"Jing K, Xu J. A survey on neural network language models. 2019; arXiv preprint arXiv:1906.03591."},{"key":"67_CR9","unstructured":"Maus N, Chao P, Wong E, Gardner J. Adversarial Prompting for Black Box Foundation Models. 2023; arXiv preprint arXiv:2302.04237."},{"key":"67_CR10","unstructured":"Sun L, Hashimoto K, Yin W, Asai A, Li J, Yu P, Xiong C. Adv-bert: Bert is not robust on misspellings! generating nature adversarial samples on bert. 2020; arXiv preprint arXiv:2003.04985."},{"issue":"6630","key":"67_CR11","doi-asserted-by":"publisher","first-page":"313","DOI":"10.1126\/science.adg7879","volume":"379","author":"HH Thorp","year":"2023","unstructured":"Thorp HH. ChatGPT is fun, but not an author. Science. 2023;379(6630):313\u2013313.","journal-title":"Science"},{"key":"67_CR12","unstructured":"Microsoft Bets Billions on DALL-E and ChatGPT Maker OpenAI. https:\/\/risnews.com\/microsoft-bets-billions-dall-e-and-chatgpt-maker-openai."},{"key":"67_CR13","unstructured":"What is generative AI? McKinsey and Company. https:\/\/www.mckinsey.com\/featured-insights\/mckinsey-explainers\/what-is-generative-ai."},{"key":"67_CR14","unstructured":"A number of papers on large language models were presented (including some in their own dedicated session) in CogSci 2022: https:\/\/cognitivesciencesociety.org\/cogsci-2022\/"},{"key":"67_CR15","unstructured":"Collins KM, Wong C, Feng J, Wei M, Tenenbaum JB. Structured, flexible, and robust: benchmarking and improving large language models towards more human-like behavior in out-of-distribution reasoning tasks. 2022; arXiv preprint arXiv:2205.05718."},{"key":"67_CR16","unstructured":"Kejriwal M, Tang Z. Evaluating language representation models on approximately rational decision making problems. In: Proceedings of the Annual Meeting of the Cognitive Science Society. 2022;44(44)."},{"issue":"11","key":"67_CR17","doi-asserted-by":"publisher","first-page":"942","DOI":"10.3357\/AMHP.4343.2015","volume":"86","author":"M Basner","year":"2015","unstructured":"Basner M, Savitt A, Moore TM, Port AM, McGuire S, Ecker AJ, et al. Development and validation of the cognition test battery for spaceflight. Aerospace Med Human Perform. 2015;86(11):942\u201352.","journal-title":"Aerospace Med Human Perform"},{"issue":"6","key":"67_CR18","doi-asserted-by":"publisher","first-page":"1003","DOI":"10.1007\/s10071-017-1135-1","volume":"20","author":"RC Shaw","year":"2017","unstructured":"Shaw RC, Schmelz M. Cognitive test batteries in animal cognition research: evaluating the past, present and future of comparative psychometrics. Anim Cogn. 2017;20(6):1003\u201318.","journal-title":"Anim Cogn"},{"issue":"11","key":"67_CR19","doi-asserted-by":"publisher","first-page":"1613","DOI":"10.1212\/01.wnl.0000434309.85312.19","volume":"55","author":"PS Mathuranath","year":"2000","unstructured":"Mathuranath PS, Nestor PJ, Berrios GE, Rakowicz W, Hodges JR. A brief cognitive test battery to differentiate Alzheimer\u2019s disease and frontotemporal dementia. Neurology. 2000;55(11):1613\u201320.","journal-title":"Neurology"},{"issue":"4","key":"67_CR20","doi-asserted-by":"publisher","first-page":"318","DOI":"10.1038\/s42256-022-00478-4","volume":"4","author":"M Kejriwal","year":"2022","unstructured":"Kejriwal M, Santos H, Mulvehill AM, McGuinness DL. Designing a strong test for measuring true common-sense reasoning. Nat Mach Intell. 2022;4(4):318\u201322.","journal-title":"Nat Mach Intell"},{"key":"67_CR21","unstructured":"Borji A. Generated faces in the wild: Quantitative comparison of stable diffusion, midjourney and dall-e 2. 2022; arXiv preprint arXiv:2210.00586."},{"key":"67_CR22","doi-asserted-by":"crossref","unstructured":"Shen K, Kejriwal M. An experimental study measuring the generalization of fine\u2010tuned language representation models across commonsense reasoning benchmarks. Expert Syst. 2023: e13243.","DOI":"10.1111\/exsy.13243"},{"key":"67_CR23","doi-asserted-by":"crossref","unstructured":"Tang Z, Kejriwal M. Can language representation models think in bets?. 2022; arXiv preprint arXiv:2210.07519.","DOI":"10.1098\/rsos.221585"},{"key":"67_CR24","unstructured":"Marcus G, Davis E, Aaronson S. A very preliminary analysis of dall-e 2. 2022; arXiv preprint arXiv:2204.13807."},{"key":"67_CR25","doi-asserted-by":"crossref","unstructured":"Leivada E, Murphy E, Marcus G. DALL-E 2 Fails to Reliably Capture Common Syntactic Processes. 2022; arXiv preprint arXiv:2210.12889.","DOI":"10.1016\/j.ssaho.2023.100648"},{"issue":"2","key":"67_CR26","doi-asserted-by":"publisher","first-page":"217","DOI":"10.1017\/S0140525X00029733","volume":"16","author":"B Landau","year":"1993","unstructured":"Landau B, Jackendoff R. \u201cWhat\u201d and \u201cwhere\u201d in spatial language and spatial cognition. Behav Brain Sci. 1993;16(2):217\u201338. https:\/\/doi.org\/10.1017\/S0140525X00029733.","journal-title":"Behav Brain Sci"},{"key":"67_CR27","doi-asserted-by":"publisher","first-page":"99","DOI":"10.2307\/1884852","volume":"69","author":"HA Simon","year":"1955","unstructured":"Simon HA. A behavioral model of rational choice. Quarter J Econ. 1955;69:99\u2013118. https:\/\/doi.org\/10.2307\/1884852.","journal-title":"Quarter J Econ"},{"key":"67_CR28","unstructured":"Von Neumann J, Morgenstern O. Theory of games and economic behavior. 1944."},{"key":"67_CR29","unstructured":"Gokhale T, Palangi H, Nushi B, Vineet V, Horvitz E, Kamar E et al. Benchmarking spatial relationships in text-to-image generation. 2022; arXiv preprint arXiv:2212.10015."},{"issue":"1","key":"67_CR30","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1515\/jci-2016-0005","volume":"4","author":"J Pearl","year":"2016","unstructured":"Pearl J. The sure-thing principle. J Causal Inference. 2016;4(1):81\u20136.","journal-title":"J Causal Inference"},{"issue":"4","key":"67_CR31","doi-asserted-by":"publisher","first-page":"967","DOI":"10.1111\/j.1744-6570.2007.00098.x","volume":"60","author":"RP Tett","year":"2007","unstructured":"Tett RP, Christiansen ND. Personality tests at the crossroads: a response to Morgeson, Campion, Dipboye, Hollenbeck, Murphy, and Schmitt (2007). Personnel Psychol. 2007;60(4):967\u201393.","journal-title":"Personnel Psychol"},{"key":"67_CR32","doi-asserted-by":"crossref","unstructured":"Bang Y, Cahyawijaya S, Lee N, Dai W, Su D, Wilie B, et al. A multitask, multilingual, multimodal evaluation of ChatGPT on reasoning, hallucination, and interactivity. 2023; arXiv preprint arXiv:2302.04023.","DOI":"10.18653\/v1\/2023.ijcnlp-main.45"}],"container-title":["Discover Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44163-023-00067-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44163-023-00067-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44163-023-00067-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,21]],"date-time":"2024-10-21T21:22:04Z","timestamp":1729545724000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44163-023-00067-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,6]]},"references-count":32,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["67"],"URL":"https:\/\/doi.org\/10.1007\/s44163-023-00067-3","relation":{},"ISSN":["2731-0809"],"issn-type":[{"value":"2731-0809","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,6]]},"assertion":[{"value":"17 February 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 May 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 June 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"21"}}