{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,28]],"date-time":"2026-02-28T03:25:01Z","timestamp":1772249101832,"version":"3.50.1"},"reference-count":83,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2022,12,21]],"date-time":"2022-12-21T00:00:00Z","timestamp":1671580800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100002283","name":"Alzheimer's Research UK","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100002283","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100000781","name":"European Research Council","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100000781","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100000265","name":"Medical Research Council","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100000265","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100010661","name":"Horizon 2020 Framework Programme","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100010661","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Comput. Neurosci."],"abstract":"<jats:sec>\n                    <jats:title>Introduction<\/jats:title>\n                    <jats:p>In recent years, machines powered by deep learning have achieved near-human levels of performance in speech recognition. The fields of artificial intelligence and cognitive neuroscience have finally reached a similar level of performance, despite their huge differences in implementation, and so deep learning models can\u2014in principle\u2014serve as candidates for mechanistic models of the human auditory system.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Methods<\/jats:title>\n                    <jats:p>Utilizing high-performance automatic speech recognition systems, and advanced non-invasive human neuroimaging technology such as magnetoencephalography and multivariate pattern-information analysis, the current study aimed to relate machine-learned representations of speech to recorded human brain representations of the same speech.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>In one direction, we found a quasi-hierarchical functional organization in human auditory cortex qualitatively matched with the hidden layers of deep artificial neural networks trained as part of an automatic speech recognizer. In the reverse direction, we modified the hidden layer organization of the artificial neural network based on neural activation patterns in human brains. The result was a substantial improvement in word recognition accuracy and learned speech representations.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Discussion<\/jats:title>\n                    <jats:p>We have demonstrated that artificial and brain neural networks can be mutually informative in the domain of speech recognition.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.3389\/fncom.2022.1057439","type":"journal-article","created":{"date-parts":[[2022,12,21]],"date-time":"2022-12-21T04:12:22Z","timestamp":1671595942000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["On the similarities of representations in artificial and brain neural networks for speech recognition"],"prefix":"10.3389","volume":"16","author":[{"given":"Cai","family":"Wingfield","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chao","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Barry","family":"Devereux","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Elisabeth","family":"Fonteneau","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrew","family":"Thwaites","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xunying","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Phil","family":"Woodland","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"William","family":"Marslen-Wilson","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Li","family":"Su","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2022,12,21]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"634","DOI":"10.1523\/JNEUROSCI.2454-14.2015","article-title":"Distributed neural representations of phonological features during speech perception","volume":"35","author":"Arsenault","year":"2015","journal-title":"J. Neurosci"},{"key":"B2","first-page":"12449","article-title":"\u201cwav2vec 2.0: a framework for self-supervised learning of speech representations,\u201d","volume-title":"Proceedings of the 34th International Conference on Neural Information Processing Systems: NIPS'20, Vol. 33","author":"Baevski","year":"2020"},{"key":"B3","doi-asserted-by":"publisher","DOI":"10.3389\/fnsys.2013.00011","article-title":"A unified framework for the organization of the primate auditory cortex","volume":"7","author":"Baumann","year":"2013","journal-title":"Front. Syst. Neurosci"},{"key":"B4","first-page":"687","article-title":"\u201cThe MGB challenge: evaluating multi-genre broadcast media transcription,\u201d","volume-title":"Proc. ASRU","author":"Bell","year":"2015"},{"key":"B5","volume-title":"Pattern Recognition and Machine Learning","author":"Bishop","year":"2006"},{"key":"B6","volume-title":"Connectionist Speech Recognition: A Hybrid Approach","author":"Bourlard","year":"1993"},{"key":"B7","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1003963","article-title":"Deep neural networks rival the representation of primate IT cortex for core visual object recognition","volume":"10","author":"Cadieu","year":"2014","journal-title":"PLOS Comput. Biol"},{"key":"B8","doi-asserted-by":"publisher","first-page":"2679","DOI":"10.1093\/cercor\/bht127","article-title":"Speech-specific tuning of neurons in human superior temporal gyrus","volume":"24","author":"Chan","year":"2014","journal-title":"Cereb. Cortex"},{"key":"B9","doi-asserted-by":"publisher","first-page":"1428","DOI":"10.1038\/nn.2641","article-title":"Categorical speech representation in human superior temporal gyrus","volume":"13","author":"Chang","year":"2010","journal-title":"Nat. Neurosci"},{"key":"B10","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2022.3188113","article-title":"WavLM: large-scale self-supervised pre-training for full stack speech processing","author":"Chen","year":"2022","journal-title":"arXiv preprint arXiv:2110.13900"},{"key":"B11","doi-asserted-by":"publisher","first-page":"346","DOI":"10.1016\/j.neuroimage.2016.03.063","article-title":"Dynamics of scene representations in the human brain revealed by magnetoencephalography and deep neural networks","volume":"153","author":"Cichy","year":"2016","journal-title":"Neuroimage"},{"key":"B12","doi-asserted-by":"publisher","first-page":"3602","DOI":"10.1093\/cercor\/bhu203","article-title":"Predicting the time course of individual objects with MEG","volume":"25","author":"Clarke","year":"2014","journal-title":"Cereb. Cortex"},{"key":"B13","doi-asserted-by":"publisher","first-page":"15015","DOI":"10.1523\/JNEUROSCI.0977-15.2015","article-title":"Decoding articulatory features from fMRI responses in dorsal speech regions","volume":"35","author":"Correia","year":"2015","journal-title":"J. Neurosci"},{"key":"B14","doi-asserted-by":"publisher","first-page":"224","DOI":"10.1109\/TPAMI.1979.4766909","article-title":"A cluster separation measure","volume":"1","author":"Davies","year":"1979","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"B15","doi-asserted-by":"publisher","first-page":"2702","DOI":"10.1121\/1.409839","article-title":"A statistical approach to automatic speech recognition using the atomic speech units constructed from overlapping articulatory features","volume":"95","author":"Deng","year":"1994","journal-title":"J. Acoust. Soc. Am"},{"key":"B16","doi-asserted-by":"publisher","first-page":"2551","DOI":"10.1523\/JNEUROSCI.3569-03.2004","article-title":"The processing of visual shape in the cerebral cortex of human and nonhuman primates: a functional magnetic resonance imaging study","volume":"24","author":"Denys","year":"2004","journal-title":"J. Neurosci"},{"key":"B17","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-018-28865-1","article-title":"Integrated deep visual and semantic attractor neural networks predict fMRI pattern-information along the ventral object processing pathway","volume":"8","author":"Devereux","year":"2018","journal-title":"Nat. Sci. Rep"},{"key":"B18","doi-asserted-by":"publisher","first-page":"2457","DOI":"10.1016\/j.cub.2015.08.030","article-title":"Low-frequency cortical entrainment to speech reflects phoneme-level processing","volume":"25","author":"Di Liberto","year":"2015","journal-title":"Curr. Biol"},{"key":"B19","doi-asserted-by":"publisher","first-page":"2199","DOI":"10.21437\/Interspeech.2014-492","article-title":"\u201cSpeaker dependent bottleneck layer training for speaker adaptation in automatic speech recognition,\u201d","author":"Doddipatla","year":"2014","journal-title":"Proc. Interspeech"},{"key":"B20","doi-asserted-by":"publisher","first-page":"3962","DOI":"10.1093\/cercor\/bhu283","article-title":"Brain network connectivity during language comprehension: interacting linguistic and perceptual subsystems","volume":"25","author":"Fonteneau","year":"2014","journal-title":"Cereb. Cortex"},{"key":"B21","doi-asserted-by":"publisher","first-page":"446","DOI":"10.1016\/j.neuroimage.2013.10.027","article-title":"MNE software for processing MEG and EEG data","volume":"86","author":"Gramfort","year":"2014","journal-title":"Neuroimage"},{"key":"B22","first-page":"757","article-title":"\u201cProbabilistic and bottle-neck features for LVCSR of meetings,\u201d","volume-title":"Proc. ICASSP","author":"Gr\u00e9zl","year":"2007"},{"key":"B23","doi-asserted-by":"publisher","first-page":"10005","DOI":"10.1523\/JNEUROSCI.5023-14.2015","article-title":"Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream","volume":"35","author":"G\u00fc\u00e7l\u00fc","year":"2015","journal-title":"J. Neurosci"},{"key":"B24","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1007\/BF02512476","article-title":"Interpreting magnetic fields of the brain: minimum norm estimates","volume":"32","author":"H\u00e4m\u00e4l\u00e4inen","year":"1994","journal-title":"Med. Biol. Eng. Comput"},{"key":"B25","doi-asserted-by":"publisher","first-page":"4626","DOI":"10.1016\/j.cell.2021.07.019","article-title":"Parallel and distributed encoding of speech across human auditory cortex","volume":"12","author":"Hamilton","year":"2021","journal-title":"Cell"},{"key":"B26","doi-asserted-by":"publisher","first-page":"82","DOI":"10.1109\/MSP.2012.2205597","article-title":"Deep neural networks for acoustic modeling in speech recognition: the shared views of four research groups","volume":"29","author":"Hinton","year":"2012","journal-title":"IEEE Signal Process. Mag"},{"key":"B27","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput"},{"key":"B28","doi-asserted-by":"publisher","first-page":"3451","DOI":"10.1109\/TASLP.2021.3122291","article-title":"HuBERT: self-supervised speech representation learning by masked prediction of hidden units","volume":"29","author":"Hsu","year":"2021","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process"},{"key":"B29","doi-asserted-by":"publisher","first-page":"782","DOI":"10.1016\/j.neuroimage.2011.09.015","article-title":"FSL","volume":"62","author":"Jenkinson","year":"2012","journal-title":"Neuroimage"},{"key":"B30","first-page":"2589","article-title":"\u201cBUT BABEL system for spontaneous Cantonese,\u201d","volume-title":"Proc. Interspeech","author":"Karafi\u00e1t","year":"2013"},{"key":"B31","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1003915","article-title":"Deep supervised, but not unsupervised, models may explain IT cortical representation","volume":"10","author":"Khaligh-Razavi","year":"2014","journal-title":"PLoS Comput. Biol"},{"key":"B32","doi-asserted-by":"publisher","DOI":"10.1038\/srep32672","article-title":"Deep networks can resemble human feed-forward vision in invariant object recognition","volume":"6","author":"Kheradpisheh","year":"2016","journal-title":"Nat. Sci. Rep"},{"key":"B33","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1146\/annurev-vision-082114-035447","article-title":"Deep neural networks: a new framework for modeling biological vision and brain information processing","volume":"1","author":"Kriegeskorte","year":"2015","journal-title":"Annu. Rev. Vision Sci"},{"key":"B34","doi-asserted-by":"publisher","first-page":"1148","DOI":"10.1038\/s41593-018-0210-5","article-title":"Cognitive computational neuroscience","volume":"21","author":"Kriegeskorte","year":"2018","journal-title":"Nat. Neurosci"},{"key":"B35","doi-asserted-by":"publisher","first-page":"3863","DOI":"10.1073\/pnas.0600244103","article-title":"Information-based functional brain mapping","volume":"103","author":"Kriegeskorte","year":"2006","journal-title":"Proc. Natl. Acad. Sci. U.S.A"},{"key":"B36","doi-asserted-by":"publisher","first-page":"401","DOI":"10.1016\/j.tics.2013.06.007","article-title":"Representational geometry: Integrating cognition, computation, and the brain","volume":"17","author":"Kriegeskorte","year":"2013","journal-title":"Trends Cogn. Sci"},{"key":"B37","doi-asserted-by":"publisher","first-page":"4","DOI":"10.3389\/neuro.06.004.2008","article-title":"Representational similarity analysis-connecting the branches of systems neuroscience","volume":"2","author":"Kriegeskorte","year":"","journal-title":"Front. Syst. Neurosci"},{"key":"B38","doi-asserted-by":"publisher","first-page":"1126","DOI":"10.1016\/j.neuron.2008.10.043","article-title":"Matching categorical object representations in inferior temporal cortex of man and monkey","volume":"60","author":"Kriegeskorte","year":"","journal-title":"Neuron"},{"key":"B39","article-title":"\u201cImagenet classification with deep convolutional neural networks,\u201d","volume-title":"Proc. NIPS","author":"Krizhevsky","year":"2012"},{"key":"B40","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"B41","doi-asserted-by":"crossref","first-page":"3145","DOI":"10.21437\/Interspeech.2015-633","article-title":"\u201cThe Cambridge university 2014 BOLT conversational telephone Mandarin Chinese LVCSR system for speech translation,\u201d","volume-title":"Proc. Interspeech","author":"Liu","year":"2015"},{"key":"B42","first-page":"231","article-title":"\u201cRWTH ASR systems for LibriSpeech: hybrid vs attention,\u201d","volume-title":"Proc. Interspeech","author":"Luscher","year":"2019"},{"key":"B43","doi-asserted-by":"publisher","first-page":"13203","DOI":"10.1073\/pnas.1614048113","article-title":"Dynamic updating of hippocampal object representations reflects new conceptual knowledge","volume":"113","author":"Mack","year":"2016","journal-title":"Proc. Natl. Acad. Sci. U.S.A"},{"key":"B44","doi-asserted-by":"publisher","first-page":"1006","DOI":"10.1126\/science.1245994","article-title":"Phonetic feature encoding in human superior temporal gyrus","volume":"343","author":"Mesgarani","year":"2014","journal-title":"Science"},{"key":"B45","doi-asserted-by":"publisher","first-page":"899","DOI":"10.1121\/1.2816572","article-title":"Phoneme representation and classification in primary auditory cortex","volume":"123","author":"Mesgarani","year":"2008","journal-title":"J. Acoust. Soc. Am"},{"key":"B46","first-page":"7145","article-title":"\u201cArticulatory trajectories for large-vocabulary speech recognition,\u201d","volume-title":"Proc. ICASSP","author":"Mitra","year":"2013"},{"key":"B47","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2014.00225","article-title":"An anatomical and functional topography of human auditory cortical areas","volume":"8","author":"Moerel","year":"2014","journal-title":"Front. Neurosci"},{"key":"B48","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2022.3207050","article-title":"Self-supervised speech representation learning: a review","author":"Mohamed","year":"2022","journal-title":"arXiv preprint arXiv:2205.10643"},{"key":"B49","doi-asserted-by":"publisher","first-page":"1069","DOI":"10.1016\/j.neuroimage.2008.05.064","article-title":"Quantification of the benefit from integrating MEG and EEG data in minimum \u21132-norm estimation","volume":"42","author":"Molins","year":"2008","journal-title":"Neuroimage"},{"key":"B50","doi-asserted-by":"publisher","first-page":"7","DOI":"10.1109\/TASL.2011.2116010","article-title":"Deep and wide: multiple layers in automatic speech recognition","volume":"20","author":"Morgan","year":"2011","journal-title":"IEEE Trans. Audio Speech Lang. Process"},{"key":"B51","doi-asserted-by":"publisher","DOI":"10.1088\/1741-2552\/aaab6f","article-title":"Real-time classification of auditory sentences using evoked cortical activity in humans","volume":"15","author":"Moses","year":"2018","journal-title":"J. Neural Eng"},{"key":"B52","doi-asserted-by":"publisher","DOI":"10.1088\/1741-2560\/13\/5\/056004","article-title":"Neural speech recognition: Continuous phoneme decoding using spatiotemporal representations of human cortical activity","volume":"13","author":"Moses","year":"2016","journal-title":"J. Neural Eng"},{"key":"B53","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1002\/hbm.1058","article-title":"Nonparametric permutation tests for functional neuroimaging: a primer with examples","volume":"15","author":"Nichols","year":"2002","journal-title":"Hum. Brain Mapp"},{"key":"B54","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1003553","article-title":"A toolbox for representational similarity analysis","volume":"10","author":"Nili","year":"2014","journal-title":"PLoS Comput. Biol"},{"key":"B55","doi-asserted-by":"publisher","first-page":"315","DOI":"10.1016\/j.tics.2004.05.009","article-title":"Comparative mapping of higher visual areas in monkeys and humans","volume":"8","author":"Orban","year":"2004","journal-title":"Trends Cogn. Sci"},{"key":"B56","first-page":"2613","article-title":"\u201cSpecAugment: a simple data augmentation method for automatic speech recognition,\u201d","volume-title":"Proc. Interspeech","author":"Park","year":"2019"},{"key":"B57","doi-asserted-by":"crossref","first-page":"3214","DOI":"10.21437\/Interspeech.2015-647","article-title":"\u201cA time delay neural network architecture for efficient modeling of long temporal contexts,\u201d","volume-title":"Proc. Interspeech","author":"Peddinti","year":"2015"},{"key":"B58","doi-asserted-by":"publisher","first-page":"718","DOI":"10.1038\/nn.2331","article-title":"Maps and streams in the auditory cortex: nonhuman primates illuminate human speech processing","volume":"12","author":"Rauschecker","year":"2009","journal-title":"Nat. Neurosci"},{"key":"B59","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/5236.001.0001","volume-title":"Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Volume 1: Foundations","author":"Rumelhart","year":"1986"},{"key":"B60","doi-asserted-by":"publisher","first-page":"42","DOI":"10.1016\/j.heares.2013.07.016","article-title":"Tonotopic mapping of human auditory cortex","volume":"307","author":"Saenz","year":"2014","journal-title":"Hear. Res"},{"key":"B61","doi-asserted-by":"publisher","first-page":"401","DOI":"10.1109\/T-C.1969.222678","article-title":"A nonlinear mapping for data structure analysis","volume":"18","author":"Sammon","year":"1969","journal-title":"IEEE Trans. Comput"},{"key":"B62","doi-asserted-by":"crossref","first-page":"132","DOI":"10.21437\/Interspeech.2017-405","article-title":"\u201cEnglish conversational telephone speech recognition by humans and machines,\u201d","volume-title":"Proc. Interspeech","author":"Saon","year":"2017"},{"key":"B63","first-page":"5149","article-title":"\u201cJapanese and Korean voice search,\u201d","volume-title":"Proc. ICASSP","author":"Schuster","year":"2012"},{"key":"B64","doi-asserted-by":"publisher","first-page":"83","DOI":"10.1016\/j.neuroimage.2008.03.061","article-title":"Threshold-free cluster enhancement: addressing problems of smoothing, threshold dependence and localisation in cluster inference","volume":"44","author":"Smith","year":"2009","journal-title":"Neuroimage"},{"key":"B65","first-page":"97","article-title":"\u201cSpatiotemporal searchlight representational similarity analysis in EMEG source space,\u201d","volume-title":"Proc. PRNI","author":"Su","year":"2012"},{"key":"B66","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2014.00368","article-title":"Mapping tonotopic organization in human temporal cortex: representational similarity analysis in EMEG source space","volume":"8","author":"Su","year":"2014","journal-title":"Front. Neurosci"},{"key":"B67","doi-asserted-by":"publisher","DOI":"10.3389\/fnins.2016.00183","article-title":"Representation of instantaneous and short-term loudness in the human cortex","volume":"10","author":"Thwaites","year":"2016","journal-title":"Front. Neurosci"},{"key":"B68","article-title":"\u201cInterpreting and improving natural-language processing (in machines) with natural language-processing (in the brain),\u201d","author":"Toneva","year":"2019","journal-title":"33rd Conference on Neural Information Processing Systems (NeurIPS 2019)"},{"key":"B69","doi-asserted-by":"publisher","first-page":"3981","DOI":"10.1523\/JNEUROSCI.23-10-03981.2003","article-title":"Neuroimaging weighs in: humans meet macaques in \u201cprimate\u201d visual cortex","volume":"23","author":"Tootell","year":"2003","journal-title":"J. Neurosci"},{"key":"B70","doi-asserted-by":"crossref","first-page":"890","DOI":"10.21437\/Interspeech.2014-223","article-title":"\u201cAcoustic modeling with deep neural networks using raw time signal for LVCSR,\u201d","volume-title":"Proc. Interspeech","author":"T\u00fcske","year":"2014"},{"key":"B71","doi-asserted-by":"publisher","first-page":"1359","DOI":"10.1016\/S0042-6989(01)00045-1","article-title":"Mapping visual cortex in monkeys and humans using surface-based atlases","volume":"41","author":"Van Essen","year":"2001","journal-title":"Vision Res"},{"key":"B72","doi-asserted-by":"publisher","first-page":"328","DOI":"10.1109\/29.21701","article-title":"Phoneme recognition using time-delay neural networks","volume":"37","author":"Waibel","year":"1989","journal-title":"IEEE Trans. Acoust. Speech Signal Process"},{"key":"B73","doi-asserted-by":"publisher","first-page":"4136","DOI":"10.1093\/cercor\/bhx268","article-title":"Neural encoding and decoding with deep learning for dynamic natural vision","volume":"28","author":"Wen","year":"2018","journal-title":"Cereb. Cortex"},{"key":"B74","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1005617","article-title":"Relating dynamic brain states to dynamic machine states: human and machine solutions to the speech recognition problem","volume":"13","author":"Wingfield","year":"2017","journal-title":"PLoS Comput. Biol"},{"key":"B75","first-page":"639","article-title":"\u201cCambridge University transcription systems for the Multi-genre Broadcast Challenge,\u201d","volume-title":"Proc. ASRU","author":"Woodland","year":"2015"},{"key":"B76","article-title":"Google's neural machine transltion system: bridging the gap between human and machine translation","author":"Wu","year":"2016","journal-title":"arXiv preprint arXiv:1609.08144"},{"key":"B77","first-page":"5255","article-title":"\u201cThe Microsoft 2016 conversational speech recognition system,\u201d","volume-title":"Proc. ICASSP","author":"Xiong","year":"2018"},{"key":"B78","volume-title":"The HTK Book (for HTK version 3.5)","author":"Young","year":"2015"},{"key":"B79","doi-asserted-by":"crossref","first-page":"307","DOI":"10.3115\/1075812.1075885","article-title":"\u201cTree-based state tying for high accuracy acoustic modelling,\u201d","volume-title":"Proc. HLT","author":"Young","year":"1994"},{"key":"B80","first-page":"185","article-title":"\u201cExtracting deep neural network bottleneck features using low-rank matrix factorization,\u201d","volume-title":"Proc. ICASSP","author":"Yu","year":"2014"},{"key":"B81","first-page":"500","article-title":"\u201cDetection-based accented speech recognition using articulatory features,\u201d","volume-title":"Proc. ASRU","author":"Zhang","year":"2011"},{"key":"B82","first-page":"3581","article-title":"\u201cA general artificial neural network extension for HTK,\u201d","volume-title":"Proc. Interspeech","author":"Zhang","year":""},{"key":"B83","first-page":"3224","article-title":"\u201cParameterised sigmoid and ReLU hidden activation functions for DNN acoustic modelling,\u201d","volume-title":"Proc. Interspeech","author":"Zhang","year":""}],"container-title":["Frontiers in Computational Neuroscience"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fncom.2022.1057439\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,10]],"date-time":"2024-10-10T16:42:02Z","timestamp":1728578522000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fncom.2022.1057439\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,21]]},"references-count":83,"alternative-id":["10.3389\/fncom.2022.1057439"],"URL":"https:\/\/doi.org\/10.3389\/fncom.2022.1057439","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/2022.06.27.497678","asserted-by":"object"}]},"ISSN":["1662-5188"],"issn-type":[{"value":"1662-5188","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,21]]},"article-number":"1057439"}}