{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,27]],"date-time":"2026-02-27T06:24:20Z","timestamp":1772173460067,"version":"3.50.1"},"update-to":[{"DOI":"10.1371\/journal.pcbi.1013334","type":"new_version","label":"New version","source":"publisher","updated":{"date-parts":[[2025,9,2]],"date-time":"2025-09-02T00:00:00Z","timestamp":1756771200000}}],"reference-count":39,"publisher":"Public Library of Science (PLoS)","issue":"8","license":[{"start":{"date-parts":[[2025,8,25]],"date-time":"2025-08-25T00:00:00Z","timestamp":1756080000000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["1R01DC021600-01"],"award-info":[{"award-number":["1R01DC021600-01"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007114","name":"Ralph W. and Grace M. Showalter Research Trust Fund","doi-asserted-by":"publisher","award":["N\/A"],"award-info":[{"award-number":["N\/A"]}],"id":[{"id":"10.13039\/100007114","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["www.ploscompbiol.org"],"crossmark-restriction":false},"short-container-title":["PLoS Comput Biol"],"abstract":"<jats:p>For static stimuli or at gross (\u223c1-s) time scales, artificial neural networks (ANNs) that have been trained on challenging engineering tasks, like image classification and automatic speech recognition, are now the best predictors of neural responses in primate visual and auditory cortex. It is, however, unknown whether this success can be extended to spiking activity at fine time scales, which are particularly relevant to audition. Here we address this question with ANNs trained on speech audio, and acute multi-electrode recordings from the auditory cortex of squirrel monkeys. We show that layers of trained ANNs can predict the spike counts of multi-units responding to speech audio and to monkey vocalizations at bin widths of 50 ms and below. For some multi-units, the ANNs explain close to all of the explainable variance\u2014much more than traditional spectrotemporal receptive fields, and more than untrained networks. Non-primary neurons tend to be more predictable by deeper layers of the ANNs, but there is much variation by neuron, which would be invisible to coarser recording modalities.<\/jats:p>","DOI":"10.1371\/journal.pcbi.1013334","type":"journal-article","created":{"date-parts":[[2025,8,25]],"date-time":"2025-08-25T17:43:05Z","timestamp":1756143785000},"page":"e1013334","update-policy":"https:\/\/doi.org\/10.1371\/journal.pcbi.corrections_policy","source":"Crossref","is-referenced-by-count":0,"title":["Deep neural networks explain spiking activity in auditory cortex"],"prefix":"10.1371","volume":"21","author":[{"given":"Bilal","family":"Ahmed","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joshua D.","family":"Downer","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Brian J.","family":"Malone","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0053-7006","authenticated-orcid":true,"given":"Joseph G.","family":"Makin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"340","published-online":{"date-parts":[[2025,8,25]]},"reference":[{"issue":"23","key":"pcbi.1013334.ref001","doi-asserted-by":"crossref","first-page":"8619","DOI":"10.1073\/pnas.1403112111","article-title":"Performance-optimized hierarchical models predict neural responses in higher visual cortex","volume":"111","author":"DLK Yamins","year":"2014","journal-title":"Proc Natl Acad Sci U S A."},{"issue":"3","key":"pcbi.1013334.ref002","article-title":"A task-optimized neural network replicates human auditory behavior, predicts brain responses, and reveals a cortical processing hierarchy","volume":"98","author":"AJE Kell","year":"2018","journal-title":"Neuron."},{"issue":"6439","key":"pcbi.1013334.ref003","doi-asserted-by":"crossref","DOI":"10.1126\/science.aav9436","article-title":"Neural population control via deep image synthesis","volume":"364","author":"P Bashivan","year":"2019","journal-title":"Science."},{"issue":"7814","key":"pcbi.1013334.ref004","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1038\/s41586-020-2350-5","article-title":"A map of object space in primate inferotemporal cortex","volume":"583","author":"P Bao","year":"2020","journal-title":"Nature."},{"issue":"12","key":"pcbi.1013334.ref005","doi-asserted-by":"crossref","first-page":"695","DOI":"10.1038\/s41583-020-00393-w","article-title":"The macaque face patch system: a turtle\u2019s underbelly for the brain","volume":"21","author":"JK Hesse","year":"2020","journal-title":"Nat Rev Neurosci."},{"issue":"1","key":"pcbi.1013334.ref006","doi-asserted-by":"crossref","first-page":"5540","DOI":"10.1038\/s41467-021-25409-6","article-title":"Computational models of category-selective brain regions enable high-throughput tests of selectivity","volume":"12","author":"NA Ratan Murty","year":"2021","journal-title":"Nat Commun."},{"key":"pcbi.1013334.ref007","unstructured":"G\u00fc\u00e7l\u00fc U, Thielen J, Hanke M, Van Gerven MAJ. Brains on beats. In: Advances in Neural Information Processing Systems. 2016. p. 2109\u201317."},{"key":"pcbi.1013334.ref008","doi-asserted-by":"crossref","unstructured":"Millet J, King JR. Inductive biases, pretraining and fine-tuning jointly account for brain responses to speech. 2021.","DOI":"10.31219\/osf.io\/fq6gd"},{"key":"pcbi.1013334.ref009","first-page":"1","article-title":"Toward a realistic model of speech processing in the brain with self-supervised learning","volume":"35","author":"J Millet","year":"2022","journal-title":"Advances in Neural Information Processing Systems."},{"key":"pcbi.1013334.ref010","unstructured":"Vaidya AR, Jain S, Huth AG. Self-supervised models of audio effectively explain human cortical responses to speech. In: Proceedings of the 39th International Conference on Machine Learning, 2022."},{"issue":"4","key":"pcbi.1013334.ref011","doi-asserted-by":"crossref","first-page":"664","DOI":"10.1038\/s41593-023-01285-9","article-title":"Intermediate acoustic-to-semantic representations link behavioral and neural responses to natural sounds","volume":"26","author":"BL Giordano","year":"2023","journal-title":"Nat Neurosci."},{"issue":"12","key":"pcbi.1013334.ref012","doi-asserted-by":"crossref","DOI":"10.1371\/journal.pbio.3002366","article-title":"Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions","volume":"21","author":"G Tuckute","year":"2023","journal-title":"PLoS Biol."},{"issue":"1","key":"pcbi.1013334.ref013","doi-asserted-by":"crossref","first-page":"148","DOI":"10.1152\/jn.00709.2020","article-title":"Temporally precise population coding of dynamic sounds by auditory cortex","volume":"126","author":"JD Downer","year":"2021","journal-title":"J Neurophysiol."},{"issue":"12","key":"pcbi.1013334.ref014","doi-asserted-by":"crossref","first-page":"2213","DOI":"10.1038\/s41593-023-01468-4","article-title":"Dissecting neural computations in the human auditory pathway using deep neural networks for speech","volume":"26","author":"Y Li","year":"2023","journal-title":"Nat Neurosci."},{"key":"pcbi.1013334.ref015","unstructured":"Collobert R, Puhrsch C, Synnaeve G. Wav2Letter: an end-to-end ConvNet-based speech recognition system. In: International Conference on Learning Representations; 2017. p. 1\u20138."},{"key":"pcbi.1013334.ref016","doi-asserted-by":"crossref","unstructured":"Wang C, Tang Y, Ma X, Wu A, Popuri S, Okhonko D, et al. FAIRSEQ S2T: fast speech-to-text modeling with fairseq. In: Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing: System Demonstrations; 2020. p. 33\u20139.","DOI":"10.18653\/v1\/2020.aacl-demo.6"},{"key":"pcbi.1013334.ref017","unstructured":"Amodei D, Ananthanarayanan S, Anubhai R, Bai J, Battenberg E, Case C. Deep-Speech 2: End-to-end Speech Recognition in English and Mandarin. In: International Conference on Machine Learning, 2016. p. 28."},{"key":"pcbi.1013334.ref018","unstructured":"Radford A, Kim JW, Xu T, Brockman G, Mcleavey C, Sutskever I. Robust speech recognition via large-scale weak supervision. In: Proceedings of the 40th International Conference on Machine Learning; 2023. p. 28492\u2013518."},{"key":"pcbi.1013334.ref019","unstructured":"Baevski A, Zhou H, Mohamed A, Auli M. Wav2vec 2.0: a framework for self-supervised learning of speech representations. In: Advances in Neural Information Processing Systems. 2020. p. 1\u201312."},{"issue":"3","key":"pcbi.1013334.ref020","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1080\/net.12.3.289.316","article-title":"Estimating spatio-temporal receptive fields of auditory and visual neurons from their responses to natural stimuli","volume":"12","author":"FE Theunissen","year":"2001","journal-title":"Network: Computation in Neural Systems."},{"issue":"5","key":"pcbi.1013334.ref021","doi-asserted-by":"crossref","DOI":"10.1371\/journal.pcbi.1011110","article-title":"A convolutional neural network provides a generalizable model of natural sound coding by neural populations in auditory cortex","volume":"19","author":"JR Pennington","year":"2023","journal-title":"PLoS Comput Biol."},{"key":"pcbi.1013334.ref022","unstructured":"Patil S. Speech2Text. 2023. https:\/\/huggingface.co\/docs\/transformers\/model_doc\/speech_to_text"},{"key":"pcbi.1013334.ref023","unstructured":"von Platen P. Wav2Vec2. 2023. https:\/\/huggingface.co\/docs\/transformers\/model_doc\/wav2vec2"},{"key":"pcbi.1013334.ref024","unstructured":"Ba JL, Kiros JR, Hinton GE. Layer normalization. 2016."},{"key":"pcbi.1013334.ref025","unstructured":"Hendrycks D, Gimpel K. Gaussian error linear units (GELUs). 2016."},{"key":"pcbi.1013334.ref026","unstructured":"Narenthiran S. Deepspeech.pytorch. 2022. https:\/\/github.com\/SeanNaren\/deepspeech.pytorch"},{"issue":"8","key":"pcbi.1013334.ref027","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"S Hochreiter","year":"1997","journal-title":"Neural Computation."},{"key":"pcbi.1013334.ref028","unstructured":"OpenAI. Whisper. 2023. https:\/\/huggingface.co\/docs\/transformers\/en\/model_doc\/whisper#whisper"},{"key":"pcbi.1013334.ref029","doi-asserted-by":"crossref","unstructured":"Mischler G, Raghavan V, Keshishian M, Mesgarani N. Naplib-Python: Neural Acoustic Data Processing and Analysis Tools in Python. 2023.","DOI":"10.1016\/j.simpa.2023.100541"},{"key":"pcbi.1013334.ref030","doi-asserted-by":"crossref","unstructured":"Karpov A, Jokisch O, Potapova R. Speech and computer. Springer; 2018. https:\/\/doi.org\/10.1007\/978-3-319-99579-3","DOI":"10.1007\/978-3-319-99579-3"},{"key":"pcbi.1013334.ref031","unstructured":"PyTorch. Torchaudio datasets: TED-LIUM. TorchAudio datasets. 2024. https:\/\/pytorch.org\/audio\/stable\/generated\/torchaudio.datasets.TEDLIUM.html#torchaudio.datasets.TEDLIUM"},{"key":"pcbi.1013334.ref032","unstructured":"Ardila R, Branson M, Davis K, Henretty M, Kohler M, Meyer J, et al. Common voice: a massively-multilingual speech corpus. LREC 2020 - 12th International Conference on Language Resources and Evaluation, Conference Proceedings. 2020; p. 4218\u201322."},{"key":"pcbi.1013334.ref033","unstructured":"Mozilla. Mozilla datasets: Common Voice Corpus 5.1. 2024. https:\/\/commonvoice.mozilla.org\/en\/datasets"},{"key":"pcbi.1013334.ref034","doi-asserted-by":"crossref","unstructured":"Wang C, Riviere M, Lee A, Wu A, Talnikar C, Haziza D, et al. VoxPopuli: a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. https:\/\/doi.org\/10.18653\/v1\/2021.acl-long.80","DOI":"10.18653\/v1\/2021.acl-long.80"},{"key":"pcbi.1013334.ref035","unstructured":"Research M. Meta research: VoxPopuli. 2024. https:\/\/github.com\/facebookresearch\/voxpopuli"},{"issue":"2","key":"pcbi.1013334.ref036","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1088\/0954-898X_15_2_002","article-title":"Quantifying variability in neural responses and its application for the validation of model predictions","volume":"15","author":"A Hsu","year":"2004","journal-title":"Network: Computation in Neural Systems."},{"key":"pcbi.1013334.ref037","doi-asserted-by":"crossref","unstructured":"Gemmeke JF, Ellis DPW, Freedman D, Jansen A, Lawrence W, Moore RC. Audio set: an ontology and human-labeled dataset for audio events. In: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 2017. p. 776\u201380.","DOI":"10.1109\/ICASSP.2017.7952261"},{"issue":"39","key":"pcbi.1013334.ref038","doi-asserted-by":"crossref","first-page":"16976","DOI":"10.1073\/pnas.1012656107","article-title":"Millisecond encoding precision of auditory cortex neurons","volume":"107","author":"C Kayser","year":"2010","journal-title":"Proc Natl Acad Sci U S A."},{"key":"pcbi.1013334.ref039","unstructured":"van Zwol B, Jefferson R, van den Broek EL. Predictive coding networks and inference learning: tutorial and survey. 2024."}],"updated-by":[{"DOI":"10.1371\/journal.pcbi.1013334","type":"new_version","label":"New version","source":"publisher","updated":{"date-parts":[[2025,9,2]],"date-time":"2025-09-02T00:00:00Z","timestamp":1756771200000}}],"container-title":["PLOS Computational Biology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dx.plos.org\/10.1371\/journal.pcbi.1013334","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,2]],"date-time":"2025-09-02T20:08:34Z","timestamp":1756843714000},"score":1,"resource":{"primary":{"URL":"https:\/\/dx.plos.org\/10.1371\/journal.pcbi.1013334"}},"subtitle":[],"editor":[{"given":"Jian","family":"Liu","sequence":"first","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2025,8,25]]},"references-count":39,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2025,8,25]]}},"URL":"https:\/\/doi.org\/10.1371\/journal.pcbi.1013334","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/2024.11.12.623280","asserted-by":"object"}]},"ISSN":["1553-7358"],"issn-type":[{"value":"1553-7358","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,25]]}}}