{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T23:16:00Z","timestamp":1776467760276,"version":"3.51.2"},"reference-count":51,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2022,9,8]],"date-time":"2022-09-08T00:00:00Z","timestamp":1662595200000},"content-version":"vor","delay-in-days":250,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,9,7]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Recent computational models of the acquisition of spoken language via grounding in perception exploit associations between spoken and visual modalities and learn to represent speech and visual data in a joint vector space. A major unresolved issue from the point of ecological validity is the training data, typically consisting of images or videos paired with spoken descriptions of what is depicted. Such a setup guarantees an unrealistically strong correlation between speech and the visual data. In the real world the coupling between the linguistic and the visual modality is loose, and often confounded by correlations with non-semantic aspects of the speech signal. Here we address this shortcoming by using a dataset based on the children\u2019s cartoon Peppa Pig. We train a simple bi-modal architecture on the portion of the data consisting of dialog between characters, and evaluate on segments containing descriptive narrations. Despite the weak and confounded signal in this training data, our model succeeds at learning aspects of the visual semantics of spoken language.<\/jats:p>","DOI":"10.1162\/tacl_a_00498","type":"journal-article","created":{"date-parts":[[2022,9,8]],"date-time":"2022-09-08T13:57:20Z","timestamp":1662645440000},"page":"922-936","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":10,"title":["Learning English with Peppa Pig"],"prefix":"10.1162","volume":"10","author":[{"given":"Mitja","family":"Nikolaus","sequence":"first","affiliation":[{"name":"Aix-Marseille University, France. mitja.nikolaus@univ-amu.fr"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Afra","family":"Alishahi","sequence":"additional","affiliation":[{"name":"Tilburg University, The Netherlands. a.alishahi@uvt.nl"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Grzegorz","family":"Chrupa\u0142a","sequence":"additional","affiliation":[{"name":"Tilburg University, The Netherlands. grzegorz@chrupala.me"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2022,9,7]]},"reference":[{"key":"2022090813571586600_bib1","doi-asserted-by":"publisher","first-page":"368","DOI":"10.18653\/v1\/K17-1037","article-title":"Encoding of phonology in a recurrent neural model of grounded speech","volume-title":"Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017)","author":"Alishahi","year":"2017"},{"issue":"5","key":"2022090813571586600_bib2","doi-asserted-by":"publisher","first-page":"505","DOI":"10.1177\/0002764204271506","article-title":"Television and very young children","volume":"48","author":"Anderson","year":"2005","journal-title":"American Behavioral Scientist"},{"key":"2022090813571586600_bib3","first-page":"12449","article-title":"Wav2vec 2.0: A framework for self-supervised learning of speech representations","volume-title":"Advances in Neural Information Processing Systems","author":"Baevski","year":"2020"},{"key":"2022090813571586600_bib4","article-title":"Word learning in the wild: What natural data can tell us","volume-title":"Proceedings of the Annual Meeting of the Cognitive Science Society","author":"Beekhuizen","year":"2013"},{"issue":"9","key":"2022090813571586600_bib5","doi-asserted-by":"publisher","first-page":"3253","DOI":"10.1073\/pnas.1113380109","article-title":"At 6\u20139 months, human infants know the meanings of many common nouns","volume":"109","author":"Bergelson","year":"2012","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2022090813571586600_bib6","first-page":"29","article-title":"Grounding spoken words in unlabeled video.","volume-title":"CVPR Workshops","author":"Boggust","year":"2019"},{"key":"2022090813571586600_bib7","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022090813571586600_bib8","doi-asserted-by":"publisher","first-page":"673","DOI":"10.1613\/jair.1.12967","article-title":"Visually grounded models of spoken language: A survey of datasets, architectures and evaluation techniques","volume":"73","author":"Chrupa\u0142a","year":"2022","journal-title":"Journal of Artificial Intelligence Research"},{"key":"2022090813571586600_bib9","doi-asserted-by":"publisher","first-page":"613","DOI":"10.18653\/v1\/P17-1057","article-title":"Representations of language in a model of visually grounded speech signal","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Chrupa\u0142a","year":"2017"},{"key":"2022090813571586600_bib10","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2022090813571586600_bib11","doi-asserted-by":"publisher","DOI":"10.7551\/mitpress\/11577.001.0001","volume-title":"Variability and Consistency in Early Language Learning: The Wordbank Project","author":"Frank","year":"2021"},{"key":"2022090813571586600_bib12","doi-asserted-by":"publisher","first-page":"219","DOI":"10.1145\/958432.958474","article-title":"A visually grounded natural language interface for reference to spatial scenes","volume-title":"Proceedings of the 5th International Conference on Multimodal Interfaces","author":"Gorniak","year":"2003"},{"key":"2022090813571586600_bib13","doi-asserted-by":"publisher","first-page":"237","DOI":"10.1109\/ASRU.2015.7404800","article-title":"Deep multimodal semantic embeddings for speech and images","volume-title":"2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU)","author":"Harwath","year":"2015"},{"key":"2022090813571586600_bib14","doi-asserted-by":"publisher","first-page":"506","DOI":"10.18653\/v1\/P17-1047","article-title":"Learning word-like units from joint audio-visual analysis","volume-title":"Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Harwath","year":"2017"},{"key":"2022090813571586600_bib15","doi-asserted-by":"publisher","first-page":"3017","DOI":"10.1109\/ICASSP.2019.8682666.","article-title":"Towards visually grounded sub-word speech unit discovery","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2019, Brighton, United Kingdom, May 12-17, 2019","author":"Harwath","year":"2019"},{"key":"2022090813571586600_bib16","doi-asserted-by":"publisher","first-page":"649","DOI":"10.1007\/978-3-030-01231-1_40","article-title":"Jointly discovering visual objects and spoken words from raw sensory input","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Harwath","year":"2018"},{"key":"2022090813571586600_bib17","first-page":"1858","article-title":"Unsupervised learning of spoken language with visual context","volume-title":"Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain","author":"Harwath","year":"2016"},{"key":"2022090813571586600_bib18","doi-asserted-by":"publisher","first-page":"8618","DOI":"10.1109\/ICASSP.2019.8683069","article-title":"Models of visually grounded speech signal pay attention to nouns: A bilingual experiment on English and Japanese","volume-title":"IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2019, Brighton, United Kingdom, May 12-17, 2019","author":"Havard","year":"2019"},{"key":"2022090813571586600_bib19","doi-asserted-by":"publisher","first-page":"pages 339\u2013pages 348","DOI":"10.18653\/v1\/K19-1032","article-title":"Word recognition, competition, and activation in a model of visually grounded speech","volume-title":"Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL)","author":"Havard","year":"2019"},{"key":"2022090813571586600_bib20","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/4575.003.0009","article-title":"The intermodal preferential looking paradigm: A window onto emerging language comprehension","volume-title":"Methods for assessing children\u2019s syntax","author":"Hirsh-Pasek","year":"1996"},{"key":"2022090813571586600_bib21","doi-asserted-by":"crossref","first-page":"3451","DOI":"10.1109\/TASLP.2021.3122291","article-title":"Hubert: Self-supervised speech representation learning by masked prediction of hidden units","volume":"29","author":"Hsu","year":"2021","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"2022090813571586600_bib22","doi-asserted-by":"publisher","first-page":"3242","DOI":"10.21437\/Interspeech.2019-1227","article-title":"Transfer learning from audio-visual grounding to speech recognition","volume-title":"Proceedings of Interspeech 2019","author":"Hsu","year":"2019"},{"key":"2022090813571586600_bib23","first-page":"1889","article-title":"Deep fragment embeddings for bidirectional image sentence mapping","volume-title":"Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems 2014, December 8-13 2014, Montreal, Quebec, Canada","author":"Karpathy","year":"2014"},{"key":"2022090813571586600_bib24","article-title":"The Kinetics human action video dataset","author":"Kay","year":"2017","journal-title":"CoRR"},{"key":"2022090813571586600_bib25","doi-asserted-by":"publisher","first-page":"123","DOI":"10.31234\/osf.io\/37zna","article-title":"Can phones, syllables, and words emerge as side- products of cross-situational audiovisual learning? A computational investigation","volume":"1","author":"Khorrami","year":"2021","journal-title":"Language Development Research"},{"key":"2022090813571586600_bib26","article-title":"Adam: A method for stochastic optimization","volume-title":"3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings","author":"Kingma","year":"2015"},{"issue":"1","key":"2022090813571586600_bib27","article-title":"Peppa Pig: An innovative way to promote formulaic language in pre-primary EFL classrooms","volume":"11","author":"Kokla","year":"2021","journal-title":"Research Papers in Language Teaching & Learning"},{"issue":"15","key":"2022090813571586600_bib28","doi-asserted-by":"publisher","first-page":"9096","DOI":"10.1073\/pnas.1532872100","article-title":"Foreign-language experience in infancy: Effects of short-term exposure and social interaction on phonetic learning","volume":"100","author":"Kuhl","year":"2003","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2022090813571586600_bib29","article-title":"Automatic generation of naturalistic child-adult interaction data","volume-title":"Proceedings of the Annual Meeting of the Cognitive Science Society","author":"Matusevych","year":"2013"},{"key":"2022090813571586600_bib30","doi-asserted-by":"publisher","first-page":"1841","DOI":"10.21437\/Interspeech.2019-3067","article-title":"Language learning using speech to image retrieval","volume-title":"Proceedings of Interspeech 2019","author":"Merkx","year":"2019"},{"key":"2022090813571586600_bib31","doi-asserted-by":"publisher","first-page":"2630","DOI":"10.1109\/ICCV.2019.00272","article-title":"Howto100m: Learning a text-video embedding by watching hundred million narrated video clips","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019","author":"Miech","year":"2019"},{"key":"2022090813571586600_bib32","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01463","article-title":"Spoken moments: Learning joint audio-visual representations from video descriptions","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Monfort","year":"2021"},{"key":"2022090813571586600_bib33","doi-asserted-by":"crossref","DOI":"10.21437\/Eurospeech.2003-635","article-title":"A visual context-aware multimodal system for spoken language processing","volume-title":"Eighth European Conference on Speech Communication and Technology","author":"Mukherjee","year":"2003"},{"key":"2022090813571586600_bib34","doi-asserted-by":"publisher","first-page":"pages 200\u2013pages 210","DOI":"10.18653\/v1\/2021.cmcl-1.24","article-title":"Evaluating the acquisition of semantic knowledge from cross-situational learning in artificial neural networks","volume-title":"Proceedings of the Workshop on Cognitive Modeling and Computational Linguistics","author":"Nikolaus","year":"2021"},{"issue":"5","key":"2022090813571586600_bib35","doi-asserted-by":"publisher","first-page":"963","DOI":"10.1111\/j.1551-6709.2011.01175.x","article-title":"Comprehension of argument structure and semantic roles: Evidence from English-learning children and the forced- choice pointing paradigm","volume":"35","author":"Noble","year":"2011","journal-title":"Cognitive Science"},{"key":"2022090813571586600_bib36","article-title":"Narration generation for cartoon videos","author":"Papasarantopoulos","year":"2021"},{"key":"2022090813571586600_bib37","first-page":"8024","article-title":"Pytorch: An imperative style, high-performance deep learning library","volume-title":"Advances in Neural Information Processing Systems 32","author":"Paszke","year":"2019"},{"key":"2022090813571586600_bib38","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9747103","article-title":"Fast-slow transformer for visually grounding speech","volume-title":"Proceedings of the 2022 International Conference on Acoustics, Speech and Signal Processing","author":"Peng","year":"2022"},{"issue":"1","key":"2022090813571586600_bib39","doi-asserted-by":"publisher","first-page":"27","DOI":"10.1348\/026151008X320156","article-title":"Just a talking book? Word learning from watching baby videos","volume":"27","author":"Robb","year":"2009","journal-title":"British Journal of Developmental Psychology"},{"key":"2022090813571586600_bib40","doi-asserted-by":"publisher","first-page":"1584","DOI":"10.21437\/Interspeech.2021-1312","article-title":"AVLnet: Learning audio-visual language representations from instructional videos","volume-title":"Proceedings of Interspeech 2021","author":"Rouditchenko","year":"2021"},{"issue":"41","key":"2022090813571586600_bib41","doi-asserted-by":"publisher","first-page":"12663","DOI":"10.1073\/pnas.1419773112","article-title":"Predicting the birth of a spoken word","volume":"112","author":"Roy","year":"2015","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2022090813571586600_bib42","unstructured":"Deb Roy . 1999. Learning from Sights and Sounds: A Computational Model. Ph.D. thesis, MIT Media Laboratory."},{"issue":"3-4","key":"2022090813571586600_bib43","doi-asserted-by":"publisher","first-page":"353","DOI":"10.1016\/S0885-2308(02)00024-4","article-title":"Learning visually grounded words and syntax for a scene description task","volume":"16","author":"Roy","year":"2002","journal-title":"Computer Speech & Language"},{"issue":"1","key":"2022090813571586600_bib44","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1207\/s15516709cog2601_4","article-title":"Learning words from sights and sounds: A computational model","volume":"26","author":"Roy","year":"2002","journal-title":"Cognitive Science"},{"issue":"5294","key":"2022090813571586600_bib45","doi-asserted-by":"publisher","first-page":"1926","DOI":"10.1126\/science.274.5294.1926","article-title":"Statistical learning by 8-month-old infants","volume":"274","author":"Saffran","year":"1996","journal-title":"Science"},{"key":"2022090813571586600_bib46","unstructured":"Thomas Schatz . 2016. ABX-Discriminability Measures and Applications. Ph.D. thesis, Universit\u00e9 Paris 6 (UPMC)."},{"issue":"1","key":"2022090813571586600_bib47","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1111\/ijal.12298","article-title":"The Peppa Pig television series as input in pre-primary EFL instruction: A corpus-based study","volume":"31","author":"Scheffler","year":"2021","journal-title":"International Journal of Applied Linguistics"},{"key":"2022090813571586600_bib48","article-title":"Learning words from images and speech","volume-title":"NIPS Workshop on Learning Semantics","author":"Synnaeve","year":"2014"},{"key":"2022090813571586600_bib49","doi-asserted-by":"publisher","first-page":"6450","DOI":"10.1109\/CVPR.2018.00675","article-title":"A closer look at spatiotemporal convolutions for action recognition","volume-title":"Proceedings of the IEEE conference on Computer Vision and Pattern Recognition","author":"Tran","year":"2018"},{"issue":"1","key":"2022090813571586600_bib50","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1017\/S0261444809990267","article-title":"D\u00e9j\u00e0 vu? A decade of research on language laboratories, television and video in language learning","volume":"43","author":"Vanderplank","year":"2010","journal-title":"Language Teaching"},{"key":"2022090813571586600_bib51","article-title":"Learning deep features for scene recognition using Places database","volume-title":"Advances in Neural Information Processing Systems","author":"Zhou","year":"2014"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00498\/2042609\/tacl_a_00498.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00498\/2042609\/tacl_a_00498.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,3]],"date-time":"2024-10-03T13:15:33Z","timestamp":1727961333000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00498\/112916\/Learning-English-with-Peppa-Pig"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022]]},"references-count":51,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00498","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022]]},"published":{"date-parts":[[2022]]}}}