{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,27]],"date-time":"2026-08-27T06:07:14Z","timestamp":1787810834282,"version":"build-2784847793"},"reference-count":62,"publisher":"MIT Press","issue":"4","license":[{"start":{"date-parts":[[2024,7,30]],"date-time":"2024-07-30T00:00:00Z","timestamp":1722297600000},"content-version":"vor","delay-in-days":211,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,12,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Large Language Models (LLMs) have been criticized for failing to connect linguistic meaning to the world\u2014for failing to solve the \u201csymbol grounding problem.\u201d Multimodal Large Language Models (MLLMs) offer a potential solution to this challenge by combining linguistic representations and processing with other modalities. However, much is still unknown about exactly how and to what degree MLLMs integrate their distinct modalities\u2014and whether the way they do so mirrors the mechanisms believed to underpin grounding in humans. In humans, it has been hypothesized that linguistic meaning is grounded through \u201cembodied simulation,\u201d the activation of sensorimotor and affective representations reflecting described experiences. Across four pre-registered studies, we adapt experimental techniques originally developed to investigate embodied simulation in human comprehenders to ask whether MLLMs are sensitive to sensorimotor features that are implied but not explicit in descriptions of an event. In Experiment 1, we find sensitivity to some features (color and shape) but not others (size, orientation, and volume). In Experiment 2, we identify likely bottlenecks to explain an MLLM\u2019s lack of sensitivity. In Experiment 3, we find that despite sensitivity to implicit sensorimotor features, MLLMs cannot fully account for human behavior on the same task. Finally, in Experiment 4, we compare the psychometric predictive power of different MLLM architectures and find that ViLT, a single-stream architecture, is more predictive of human responses to one sensorimotor feature (shape) than CLIP, a dual-encoder architecture\u2014despite being trained on orders of magnitude less data. These results reveal strengths and limitations in the ability of current MLLMs to integrate language with other modalities, and also shed light on the likely mechanisms underlying human language comprehension.<\/jats:p>","DOI":"10.1162\/coli_a_00531","type":"journal-article","created":{"date-parts":[[2024,7,30]],"date-time":"2024-07-30T15:08:33Z","timestamp":1722352113000},"page":"1415-1440","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":6,"title":["Do Multimodal Large Language Models and Humans Ground Language Similarly?"],"prefix":"10.1162","volume":"50","author":[{"given":"Cameron R.","family":"Jones","sequence":"first","affiliation":[{"name":"University of California, San Diego, Department of Cognitive Science. cameron@ucsd.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Benjamin","family":"Bergen","sequence":"additional","affiliation":[{"name":"University of California, San Diego, Department of Cognitive Science. bkbergen@ucsd.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sean","family":"Trott","sequence":"additional","affiliation":[{"name":"University of California, San Diego, Department of Cognitive Science. sttrott@ucsd.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,12,1]]},"reference":[{"issue":"4","key":"2024122021045120400_bib1","doi-asserted-by":"publisher","first-page":"577","DOI":"10.1017\/S0140525X99002149","article-title":"Perceptual symbol systems","volume":"22","author":"Barsalou","year":"1999","journal-title":"Behavioral and Brain Sciences"},{"key":"2024122021045120400_bib2","doi-asserted-by":"publisher","first-page":"5185","DOI":"10.18653\/v1\/2020.acl-main.463","article-title":"Climbing towards NLU: On meaning, form, and understanding in the age of data","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Bender","year":"2020"},{"key":"2024122021045120400_bib3","first-page":"142","article-title":"Embodiment, simulation and meaning","volume-title":"The Routledge Handbook of Semantics","author":"Bergen","year":"2015"},{"key":"2024122021045120400_bib4","volume-title":"Louder than Words: The New Science of How the Mind Makes Meaning","author":"Bergen","year":"2012"},{"issue":"11","key":"2024122021045120400_bib5","doi-asserted-by":"publisher","first-page":"527","DOI":"10.1016\/j.tics.2011.10.001","article-title":"The neurobiology of semantic memory","volume":"15","author":"Binder","year":"2011","journal-title":"Trends in Cognitive Sciences"},{"key":"2024122021045120400_bib6","doi-asserted-by":"publisher","first-page":"8718","DOI":"10.18653\/v1\/2020.emnlp-main.703","article-title":"Experience grounds language","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Bisk","year":"2020"},{"key":"2024122021045120400_bib7","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1613\/jair.4135","article-title":"Multimodal distributional semantics","volume":"49","author":"Bruni","year":"2014","journal-title":"Journal of Artificial Intelligence Research"},{"key":"2024122021045120400_bib8","doi-asserted-by":"publisher","first-page":"978","DOI":"10.1162\/tacl_a_00408","article-title":"Multimodal pretraining unmasked: A meta-analysis and a unified framework of vision-and-language BERTs","volume":"9","author":"Bugliarello","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"issue":"2","key":"2024122021045120400_bib9","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1007\/b97636","article-title":"Multimodel inference: Understanding AIC and BIC in model selection","volume":"33","author":"Burnham","year":"2004","journal-title":"Sociological Methods & Research"},{"issue":"1","key":"2024122021045120400_bib10","doi-asserted-by":"publisher","first-page":"293","DOI":"10.1162\/coli_a_00492","article-title":"Language model behavior: A comprehensive survey","volume":"50","author":"Chang","year":"2024","journal-title":"Computational Linguistics"},{"key":"2024122021045120400_bib11","doi-asserted-by":"publisher","first-page":"112","DOI":"10.3115\/v1\/P15-2019","article-title":"Learning language through pictures","volume-title":"Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)","author":"Chrupa\u0142a","year":"2015"},{"issue":"3","key":"2024122021045120400_bib12","doi-asserted-by":"publisher","first-page":"476","DOI":"10.1016\/j.cognition.2006.02.009","article-title":"Representing object colour in language comprehension","volume":"102","author":"Connell","year":"2007","journal-title":"Cognition"},{"issue":"5","key":"2024122021045120400_bib13","doi-asserted-by":"publisher","first-page":"1097","DOI":"10.1162\/089976698300017368","article-title":"Category learning through multimodality sensing","volume":"10","author":"De Sa","year":"1998","journal-title":"Neural Computation"},{"key":"2024122021045120400_bib14","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"issue":"7","key":"2024122021045120400_bib15","doi-asserted-by":"publisher","first-page":"597","DOI":"10.1016\/j.tics.2023.04.008","article-title":"Can AI language models replace human participants?","volume":"27","author":"Dillion","year":"2023","journal-title":"Trends in Cognitive Sciences"},{"key":"2024122021045120400_bib16","article-title":"An image is worth 16x16 words: Transformers for image recognition at scale","author":"Dosovitskiy","year":"2021","journal-title":"arXiv preprint arXiv:2010.11929"},{"key":"2024122021045120400_bib17","article-title":"Palm-e: An embodied multimodal language model","author":"Driess","year":"2023","journal-title":"arXiv preprint arXiv:2303.03378"},{"issue":"2","key":"2024122021045120400_bib18","doi-asserted-by":"publisher","first-page":"156","DOI":"10.1177\/2515245919847202","article-title":"Evaluating effect size in psychological research: Sense and nonsense","volume":"2","author":"Funder","year":"2019","journal-title":"Advances in Methods and Practices in Psychological Science"},{"issue":"3\u20134","key":"2024122021045120400_bib19","doi-asserted-by":"publisher","first-page":"455","DOI":"10.1080\/02643290442000310","article-title":"The brain\u2019s concepts: The role of the sensory-motor system in conceptual knowledge","volume":"22","author":"Gallese","year":"2005","journal-title":"Cognitive Neuropsychology"},{"issue":"4","key":"2024122021045120400_bib20","doi-asserted-by":"publisher","first-page":"328","DOI":"10.1198\/000313006X152649","article-title":"The difference between \u201csignificant\u201d and \u201cnot significant\u201d is not itself statistically significant","volume":"60","author":"Gelman","year":"2006","journal-title":"The American Statistician"},{"key":"2024122021045120400_bib21","doi-asserted-by":"publisher","first-page":"15180","DOI":"10.1109\/CVPR52729.2023.01457","article-title":"ImageBind: One embedding space to bind them all","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Girdhar","year":"2023"},{"key":"2024122021045120400_bib22","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2021-698","article-title":"AST: Audio Spectrogram Transformer","author":"Gong","year":"2021","journal-title":"arXiv preprint arXiv:2104.01778"},{"issue":"1\u20133","key":"2024122021045120400_bib23","doi-asserted-by":"publisher","first-page":"335","DOI":"10.1016\/0167-2789(90)90087-6","article-title":"The symbol grounding problem","volume":"42","author":"Harnad","year":"1990","journal-title":"Physica D: Nonlinear Phenomena"},{"key":"2024122021045120400_bib24","doi-asserted-by":"publisher","first-page":"237","DOI":"10.1109\/ASRU.2015.7404800","article-title":"Deep multimodal semantic embeddings for speech and images","volume-title":"2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU)","author":"Harwath","year":"2015"},{"issue":"2","key":"2024122021045120400_bib25","doi-asserted-by":"publisher","first-page":"301","DOI":"10.1016\/S0896-6273(03)00838-9","article-title":"Somatotopic representation of action words in human motor and premotor cortex","volume":"41","author":"Hauk","year":"2004","journal-title":"Neuron"},{"key":"2024122021045120400_bib26","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.230","article-title":"A fine-grained comparison of pragmatic language understanding in humans and language models","author":"Hu","year":"2022","journal-title":"arXiv preprint arXiv:2212.06801"},{"key":"2024122021045120400_bib27","article-title":"Language is not all you need: Aligning perception with language models","author":"Huang","year":"2023","journal-title":"arXiv preprint arXiv:2302.14045"},{"key":"2024122021045120400_bib28","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.5143773","article-title":"OpenCLIP (0.1)","author":"Ilharco","year":"2021","journal-title":"Zenodo"},{"key":"2024122021045120400_bib29","first-page":"482","article-title":"Distributional semantics still can\u2019t account for affordances","volume-title":"Proceedings of the Annual Meeting of the Cognitive Science Society","author":"Jones","year":"2022"},{"issue":"4","key":"2024122021045120400_bib30","doi-asserted-by":"publisher","first-page":"761","DOI":"10.1162\/COLI_a_00300","article-title":"Representation of linguistic form and function in recurrent neural networks","volume":"43","author":"K\u00e1d\u00e1r","year":"2017","journal-title":"Computational Linguistics"},{"key":"2024122021045120400_bib31","doi-asserted-by":"publisher","first-page":"3530","DOI":"10.1109\/ICCV48922.2021.00351","article-title":"Towards rotation invariance in object detection","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Kalra","year":"2021"},{"key":"2024122021045120400_bib32","doi-asserted-by":"publisher","first-page":"4933","DOI":"10.18653\/v1\/2023.emnlp-main.301","article-title":"Text encoders bottleneck compositionality in contrastive vision-language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Kamath","year":"2023"},{"key":"2024122021045120400_bib33","first-page":"5583","article-title":"ViLT: Vision-and-language transformer without convolution or region supervision","volume-title":"International Conference on Machine Learning","author":"Kim","year":"2021"},{"key":"2024122021045120400_bib34","doi-asserted-by":"publisher","first-page":"922","DOI":"10.18653\/v1\/P18-1085","article-title":"Illustrative language understanding: Large-scale visual grounding with image search","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Kiros","year":"2018"},{"key":"2024122021045120400_bib35","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-naacl.129","article-title":"Psychometric predictive power of large language models","author":"Kuribayashi","year":"2023","journal-title":"arXiv preprint arXiv:2311.07484"},{"key":"2024122021045120400_bib36","doi-asserted-by":"publisher","first-page":"740","DOI":"10.1007\/978-3-319-10602-1_48","article-title":"Microsoft COCO: Common objects in context","volume-title":"European Conference on Computer Vision","author":"Lin","year":"2014"},{"key":"2024122021045120400_bib37","doi-asserted-by":"publisher","first-page":"635","DOI":"10.1162\/tacl_a_00566","article-title":"Visual spatial reasoning","volume":"11","author":"Liu","year":"2023","journal-title":"Transactions of the Association for Computational Linguistics"},{"issue":"1\u20133","key":"2024122021045120400_bib38","doi-asserted-by":"publisher","first-page":"59","DOI":"10.1016\/j.jphysparis.2008.03.004","article-title":"A critical look at the embodied cognition hypothesis and a new proposal for grounding conceptual content","volume":"102","author":"Mahon","year":"2008","journal-title":"Journal of Physiology-Paris"},{"issue":"7","key":"2024122021045120400_bib39","doi-asserted-by":"publisher","first-page":"788","DOI":"10.1016\/j.cortex.2010.11.002","article-title":"Coming of age: A review of embodiment and the neuroscience of semantics","volume":"48","author":"Meteyard","year":"2012","journal-title":"Cortex"},{"key":"2024122021045120400_bib40","unstructured":"ML Foundations. 2023. OpenCLIP. https:\/\/github.com\/mlfoundations\/open_clip. Python package version 2.23.0."},{"key":"2024122021045120400_bib41","article-title":"The vector grounding problem","author":"Mollo","year":"2023","journal-title":"arXiv preprint arXiv:2304.01481"},{"issue":"1","key":"2024122021045120400_bib42","doi-asserted-by":"publisher","first-page":"5","DOI":"10.5334\/joc.139","article-title":"Towards strong inference in research on embodiment\u2013possibilities and limitations of causal paradigms","volume":"4","author":"Ostarek","year":"2021","journal-title":"Journal of Cognition"},{"issue":"12","key":"2024122021045120400_bib43","doi-asserted-by":"publisher","first-page":"976","DOI":"10.1038\/nrn2277","article-title":"Where do you know what you know? The representation of semantic knowledge in the human brain","volume":"8","author":"Patterson","year":"2007","journal-title":"Nature Reviews Neuroscience"},{"issue":"6","key":"2024122021045120400_bib44","doi-asserted-by":"publisher","first-page":"1108","DOI":"10.1080\/17470210802633255","article-title":"Short article: Language comprehenders retain implied shape and orientation of objects","volume":"62","author":"Pecher","year":"2009","journal-title":"Quarterly Journal of Experimental Psychology"},{"key":"2024122021045120400_bib45","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2022-10652","article-title":"Word discovery in visually grounded, self-supervised speech models","author":"Peng","year":"2022","journal-title":"arXiv preprint arXiv:2203.15081"},{"issue":"9","key":"2024122021045120400_bib46","doi-asserted-by":"publisher","first-page":"458","DOI":"10.1016\/j.tics.2013.06.004","article-title":"How neurons make meaning: Brain mechanisms for embodied and abstract-symbolic semantics","volume":"17","author":"Pulverm\u00fcller","year":"2013","journal-title":"Trends in Cognitive Sciences"},{"key":"2024122021045120400_bib47","first-page":"8748","article-title":"Learning transferable visual models from natural language supervision","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Radford","year":"2021"},{"issue":"8","key":"2024122021045120400_bib48","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford","year":"2019","journal-title":"OpenAI Blog"},{"key":"2024122021045120400_bib49","first-page":"25278","article-title":"LAION-5B: An open large-scale dataset for training next generation image-text models","volume":"35","author":"Schuhmann","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"2","key":"2024122021045120400_bib50","doi-asserted-by":"publisher","first-page":"68","DOI":"10.1145\/3624724","article-title":"Talking about large language models","volume":"67","author":"Shanahan","year":"2023","journal-title":"Communications of the ACM"},{"issue":"2","key":"2024122021045120400_bib51","doi-asserted-by":"publisher","first-page":"153","DOI":"10.1111\/1467-9280.00326","article-title":"The effect of implied orientation derived from verbal context on picture recognition","volume":"12","author":"Stanfield","year":"2001","journal-title":"Psychological Science"},{"key":"2024122021045120400_bib52","unstructured":"The HuggingFace Team and Contributors. 2023. Transformers: State-of-the-art machine learning for JAX, PyTorch and TensorFlow. https:\/\/github.com\/huggingface\/transformers. Python package version 4.35.2."},{"key":"2024122021045120400_bib53","first-page":"9568","article-title":"Eyes wide shut? Exploring the visual shortcomings of multimodal LLMs","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Tong","year":"2024"},{"issue":"5","key":"2024122021045120400_bib54","doi-asserted-by":"publisher","first-page":"1239","DOI":"10.1037\/rev0000420","article-title":"Word meaning is both categorical and continuous","volume":"30","author":"Trott","year":"2023","journal-title":"Psychological Review"},{"issue":"7","key":"2024122021045120400_bib55","doi-asserted-by":"publisher","first-page":"e13309","DOI":"10.1111\/cogs.13309","article-title":"Do large language models know what humans know?","volume":"47","author":"Trott","year":"2023","journal-title":"Cognitive Science"},{"key":"2024122021045120400_bib56","doi-asserted-by":"publisher","DOI":"10.7551\/mitpress\/9780262529365.001.0001","volume-title":"The Embodied Mind, revised edition: Cognitive Science and Human Experience","author":"Varela","year":"2017"},{"key":"2024122021045120400_bib57","first-page":"5998","article-title":"Attention is all you need","volume-title":"Advances in Neural Information Processing Systems","author":"Vaswani","year":"2017"},{"issue":"1","key":"2024122021045120400_bib58","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1515\/langcog-2012-0001","article-title":"Language comprehenders represent object distance both visually and auditorily","volume":"4","author":"Winter","year":"2012","journal-title":"Language and Cognition"},{"key":"2024122021045120400_bib59","doi-asserted-by":"publisher","first-page":"2247","DOI":"10.1109\/BigData59044.2023.10386743","article-title":"Multimodal large language models: A survey","volume-title":"2023 IEEE International Conference on Big Data","author":"Wu","year":"2023"},{"key":"2024122021045120400_bib60","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i9.26263","article-title":"BridgeTower: Building bridges between encoders in vision-language representation learning","author":"Xu","year":"2023"},{"issue":"12","key":"2024122021045120400_bib61","doi-asserted-by":"publisher","first-page":"e51382","DOI":"10.1371\/journal.pone.0051382","article-title":"Revisiting mental simulation in language comprehension: Six replication attempts","volume":"7","author":"Zwaan","year":"2012","journal-title":"PloS ONE"},{"issue":"2","key":"2024122021045120400_bib62","doi-asserted-by":"publisher","first-page":"168","DOI":"10.1111\/1467-9280.00430","article-title":"Language comprehenders mentally represent the shapes of objects","volume":"13","author":"Zwaan","year":"2002","journal-title":"Psychological Science"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/50\/4\/1415\/2470046\/coli_a_00531.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/50\/4\/1415\/2470046\/coli_a_00531.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T21:05:06Z","timestamp":1734728706000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/50\/4\/1415\/123786\/Do-Multimodal-Large-Language-Models-and-Humans"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":62,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2024,12,1]]},"published-print":{"date-parts":[[2024,12,1]]}},"URL":"https:\/\/doi.org\/10.1162\/coli_a_00531","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"value":"0891-2017","type":"print"},{"value":"1530-9312","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}