{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T04:15:44Z","timestamp":1783311344434,"version":"3.54.6"},"reference-count":73,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T00:00:00Z","timestamp":1783296000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T00:00:00Z","timestamp":1783296000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001774","name":"The University of Sydney","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001774","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Lang Resources &amp; Evaluation"],"published-print":{"date-parts":[[2026,9]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Child speech corpora are limited in number and scope, with none available for Australian English (AusE) until now, primarily due to orthographic transcription costs. Therefore, we developed AusKidTalk, the first AusE child speech corpus, a novel population due to speaker age and accent. Annotating AusKidTalk presented a circular problem: eliminating costly manual transcription required automatic speech recognition (ASR) tools not yet developed; but developing ASR tools required annotated speech corpora not available. This paper demonstrates how orthographic transcription burden was reduced in AusKidTalk via strategic data collection combined with out-of-domain ASR tools for automatic annotation augmented by manual correction. 620 children completed a single word production task, orthographic transcriptions were generated for 454 (73%) using the semi-automatic AusKidTalk pipeline and corrected manually for 394 using a custom Praat interface. An additional 58 children were transcribed without ASR assistance for comparison, yielding a total of 461 (74%) transcriptions. Workflow efficiency was evaluated on 380 children, producing a total of 136.6\u00a0h of speech. Annotation burden was reduced by (i) automatic orthographic transcription with 20% word error rate, removing the need to transcribe speech from scratch, and by (ii) the Praat interface allowing 136.6\u00a0h of speech to be corrected by listening to 18.5\u00a0h. Annotator reliability was good, with annotators achieving high similarity on 43\/49 ground truth files. Correction time for one child was typically 1.5\u00a0h compared to the 4\u00a0h for transcription unassisted by ASR. The workflow can be adapted for other corpora and updated with new ASR tools.<\/jats:p>","DOI":"10.1007\/s10579-026-09929-5","type":"journal-article","created":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T03:51:23Z","timestamp":1783309883000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["AusKidTalk: developing an orthographic annotation workflow for a speech corpus of Australian English-speaking children"],"prefix":"10.1007","volume":"60","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3879-7361","authenticated-orcid":false,"given":"T\u00fcnde","family":"Szalay","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mostafa","family":"Shahin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tharamkulasingam","family":"Sirojan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zheng","family":"Nan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Renata","family":"Huang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Joanne","family":"Arciuli","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Elise","family":"Baker","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Felicity","family":"Cox","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kirrie","family":"Ballard","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Beena","family":"Ahmed","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,7,6]]},"reference":[{"key":"9929_CR1","doi-asserted-by":"publisher","unstructured":"Ahmed, B., Ballard, K.J., Burnham, D., Tharmakulasingam, S., Mehmood, H., Estival, D. & Ambikairajah, E. (2021). AusKidTalk: An auditory-visual corpus of 3-to 12-year-old Australian children\u2019s speech. Proc. of INTERSPEECH (pp. 3680\u20133684). https:\/\/doi.org\/10.21437\/Interspeech.2021-2000","DOI":"10.21437\/Interspeech.2021-2000"},{"key":"9929_CR2","doi-asserted-by":"publisher","unstructured":"Arciuli, J., & Ballard, K. J. (2017). Still not adult-like: Lexical stress contrastivity in word productions of eight-to eleven-year-olds. Journal of Child Language, 44(5), 1274\u20131288. https:\/\/doi.org\/10.1017\/S0305000916000489","DOI":"10.1017\/S0305000916000489"},{"key":"9929_CR3","unstructured":"Arciuli, J., Ballard, K. J., Phillips, K., & Vogel, A. (2019). Using MAUS to investigate children\u2019s production of lexical stress. Proc. of the 19th International Congress of Phonetic Sciences (pp. 2470\u20132474)"},{"key":"9929_CR4","doi-asserted-by":"publisher","unstructured":"ARDC. (2022). Publishing sensitive data. Zenodo. https:\/\/doi.org\/10.5281\/zenodo.7259742","DOI":"10.5281\/zenodo.7259742"},{"key":"9929_CR5","unstructured":"Baker, E., Cox, F., Arciuli, J., & Ballard, K. J. (2019). Speech test for Australian children (STAC)"},{"key":"9929_CR6","doi-asserted-by":"crossref","unstructured":"Ballard, K. J., Djaja, D., Arciuli, J., James, D. G., & van Doorn, J. (2012). Developmental trajectory for production of prosody: Lexical stress contrastivity in children ages 3 to 7 years and in adults. Journal of Speech, Language, and Hearing Research, 55(6), 1822\u20131835.","DOI":"10.1044\/1092-4388(2012\/11-0257)"},{"key":"9929_CR7","doi-asserted-by":"crossref","unstructured":"Ballard, K. J., Robin, D. A., McCabe, P., & McDonald, J. (2010). A treatment for dysprosody in childhood apraxia of speech. Journal of Speech, Language, and Hearing Research, 53(5), 1227\u20131245.","DOI":"10.1044\/1092-4388(2010\/09-0130)"},{"key":"9929_CR8","unstructured":"Ballard, K. J., Shahin, M., & Ahmed, B. (2025). AI-driven speech therapy tool for assisting clinicians in diagnosis and therapy. Digital Health Week, 4(1)."},{"key":"9929_CR9","doi-asserted-by":"crossref","unstructured":"Bates, D., M\u00e4chler, M., Bolker, B., & Walker, S. (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1\u201348.","DOI":"10.18637\/jss.v067.i01"},{"key":"9929_CR10","doi-asserted-by":"crossref","unstructured":"Batliner, A., Blomberg, M., D\u2019Arcy, S., Elenius, D., Giuliani, D., Gerosa, M. & Wong, M. (2005). The PF_STAR children\u2019s speech corpus. Proc. of the $$9^{th}$$ European Conference on Speech Communication and Technology (pp. 2761\u20132764). Lisbon: ISCA.","DOI":"10.21437\/Interspeech.2005-705"},{"key":"9929_CR11","doi-asserted-by":"crossref","unstructured":"Bazillon, T., Esteve, Y., & Luzzati, D. (2008). Manual vs assisted transcription of prepared and spontaneous speech. In Proc. of International Conference on the Sixth Language Resources and Evaluation (pp. 1067\u20131071).","DOI":"10.63317\/4e86jjmacnuo"},{"key":"9929_CR12","doi-asserted-by":"crossref","unstructured":"Besacier, L., Barnard, E., Karpov, A., & Schultz, T. (2014). Automatic speech recognition for under-resourced languages: A survey. Speech Communication, 56, 85\u2013100.","DOI":"10.1016\/j.specom.2013.07.008"},{"key":"9929_CR13","doi-asserted-by":"crossref","unstructured":"Bhardwaj, V., Ben Othman, M. T., Kukreja, V., Belkhier, Y., Bajaj, M., Goud, B. S., & Hamam, H. (2022). Automatic speech recognition (ASR) systems for children: A systematic literature review. Applied Sciences, 12(9), 4419.","DOI":"10.3390\/app12094419"},{"key":"9929_CR14","doi-asserted-by":"crossref","unstructured":"Bhardwaj, V., Kukreja, V., & Singh, A. (2021). Usage of prosody modification and acoustic adaptation for robust automatic speech recognition (ASR) system. Revue d\u2019Intelligence Artificielle, 35(3), 235\u2013242.","DOI":"10.18280\/ria.350307"},{"key":"9929_CR15","unstructured":"Boersma, P., & Weenink, D. (2023). Praat 6.3.17. http:\/\/www.fon.hum.uva.nl\/praat\/"},{"key":"9929_CR16","doi-asserted-by":"crossref","unstructured":"Brooks, M. E., Kristensen, K., van Benthem, K. J., Magnusson, A., Berg, C. W., Nielsen, A., & Bolker, B. M. (2017). glmmTMB balances speed and flexibility among packages for zero-inflated generalized linear mixed modeling. The R Journal, 9(2), 378\u2013400.","DOI":"10.32614\/RJ-2017-066"},{"key":"9929_CR17","doi-asserted-by":"crossref","unstructured":"Burnham, D., Estival, D., Fazio, S., Viethen, J., Cox, F., Dale, R. & Hajek, J. (2011). Building an audio-visual corpus of Australian English: large corpus collection with an economical portable and replicable black box.  Proc. of INTERSPEECH ( pp. 848-851).","DOI":"10.21437\/Interspeech.2011-309"},{"key":"9929_CR18","doi-asserted-by":"crossref","unstructured":"Buttigieg, L., Grech, H., Fabri, S. G., Attard, J., & Farrugia, P. (2021). Automatic speech recognition in the assessment of child speech. In Manual of Clinical Phonetics (pp. 508\u2013515). Milton Park: Routledge.","DOI":"10.4324\/9780429320903-35"},{"key":"9929_CR19","doi-asserted-by":"crossref","unstructured":"Chen, N. F., Tong, R., Wee, D., Lee, P. X., Ma, B., & Li, H. (2016). SingaKids-Mandarin: Speech corpus of Singaporean children speaking Mandarin Chinese. In Proc. INTERSPEECH  (pp. 1545\u20131549).","DOI":"10.21437\/Interspeech.2016-139"},{"key":"9929_CR20","unstructured":"Cole, R., & Pellom, B. (2006). University of Colorado read and summarized story corpus. University of Colorado. TR-CSLR-2006-03 Technical report."},{"issue":"2\u20133","key":"9929_CR21","doi-asserted-by":"publisher","first-page":"200","DOI":"10.1080\/07268602.2024.2380680","volume":"44","author":"F Cox","year":"2024","unstructured":"Cox, F., & Penney, J. (2024). Multicultural Australian English-the new voice of Sydney. Australian Journal of Linguistics, 44(2\u20133), 200\u2013219.","journal-title":"Australian Journal of Linguistics"},{"key":"9929_CR22","unstructured":"Eskenazi, M., Mostow, J., & Graff, D. (1997). The CMU kids corpus LDC97S63. Web Download. Philadelphia: Linguistic Data Consortium"},{"key":"9929_CR23","doi-asserted-by":"crossref","unstructured":"Gao, J., Li, A. & Xiong, Z. (2012). Mandarin multimedia child speech corpus: Cass_Child.   In Proc. of International Conference on Speech Database and Assessments (pp. 7\u201312).","DOI":"10.1109\/ICSDA.2012.6422462"},{"key":"9929_CR24","doi-asserted-by":"crossref","unstructured":"Garofolo, J.S., Lamel, L.F., Fisher, W.M., Fiscus, J.G. & Pallett, D.S. (1993). DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1. NASA STI\/Recon technical report n, 93, 27403","DOI":"10.6028\/NIST.IR.4930"},{"key":"9929_CR25","unstructured":"Gibson, A., Cox, F. & Penney, J. (2023). Acquiring allophony: GOOSE and SCHOOL vowels in the speech of Australian children. In Proc. of International Congress of Phonetic Sciences (pp. 3750\u20133754 )."},{"issue":"1","key":"9929_CR26","doi-asserted-by":"publisher","first-page":"20190058","DOI":"10.1515\/lingvan-2019-0058","volume":"6","author":"S Gonzalez","year":"2020","unstructured":"Gonzalez, S., Grama, J., & Travis, C. E. (2020). Comparing the performance of forced aligners used in sociophonetic research. Linguistics Vanguard, 6(1), 20190058.","journal-title":"Linguistics Vanguard"},{"key":"9929_CR27","unstructured":"Gretter, R., Matassoni, M., Bann\u00f2, S., & Falavigna, D. (2020). TLT-school: a corpus of non native children speech arXiv:2001.08051 arXiv preprint."},{"key":"9929_CR28","doi-asserted-by":"crossref","unstructured":"H\u00e4m\u00e4l\u00e4inen, A., Cho, H., Candeias, S., Pellegrini, T., Abad, A., Tjalve, M., & Dias, M. S. (2014). Automatically recognising European Portuguese children\u2019s speech: Pronunciation patterns revealed by an analysis of ASR errors. In Computational processing of the Portuguese Language (pp. 1\u201311).","DOI":"10.1007\/978-3-319-09761-9_1"},{"key":"9929_CR29","doi-asserted-by":"crossref","unstructured":"Helgad\u00f3ttir, I. R., Kjaran, R., Nikul\u00e1sd\u00f3ttir, A. B., & Gu\u00f0nason, J. (2017). Building an ASR corpus using Althingi\u2019s parliamentary speeches. In Proc. of INTERSPEECH (pp. 2163\u20132167).","DOI":"10.21437\/Interspeech.2017-903"},{"key":"9929_CR30","unstructured":"Hu, J., Szalay, T., & Ballard, K. J. (2026). Investigating the feasibility of a novel automated speech analysis application for assessing Australian children at risk of speech sound disorders. Abstract and Talk. Speech Pathology Australia Conference. Gold Coast, Queensland, Australia"},{"key":"9929_CR31","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2021.114591","volume":"171","author":"R Jahangir","year":"2021","unstructured":"Jahangir, R., Teh, Y. W., Nweke, H. F., Mujtaba, G., Al-Garadi, M. A., & Ali, I. (2021). Speaker identification through artificial intelligence techniques: A comprehensive review and research challenges. Expert Systems with Applications, 171, Article 114591.","journal-title":"Expert Systems with Applications"},{"key":"9929_CR32","doi-asserted-by":"crossref","unstructured":"Kathania, H. K., Kadiri, S. R., Alku, P., & Kurimo, M. (2021). Using data augmentation and time-scale modification to improve ASR of children\u2019s speech in noisy environments. Applied Sciences, 11(18), 8420.","DOI":"10.3390\/app11188420"},{"key":"9929_CR33","doi-asserted-by":"crossref","unstructured":"Kathania, H. K., Kadiri, S. R., Alku, P., & Kurimo, M. (2022). A formant modification method for improved ASR of children\u2019s speech. Speech Communication, 136, 98\u2013106.","DOI":"10.1016\/j.specom.2021.11.003"},{"key":"9929_CR34","doi-asserted-by":"crossref","unstructured":"Kazemzadeh, A., You, H., Iseli, M., Jones, B., Cui, X., Heritage, M., & Alwan, A. (2005). TBALL data collection: the making of a young children\u2019s speech corpus. In Proc. of INTERSPEECH  (pp. 1581\u20131584).","DOI":"10.21437\/Interspeech.2005-462"},{"key":"9929_CR35","doi-asserted-by":"crossref","unstructured":"Kempe, V., Brooks, P. J., & Gillis, S. (2024). Four decades of open language science: The CHILDES project. Language Teaching Research Quarterly, 44, 15\u201330.","DOI":"10.32038\/ltrq.2024.44.04"},{"key":"9929_CR36","unstructured":"Kulebi, B., Armentano-Oller, C., Rodriguez-Penagos, C., & Villegas, M. (2022). ParlamentParla: A speech corpus of Catalan parliamentary sessions. In Proc. of the workshop ParlaCLARIN III within the 13tth Language Resources and Evaluation Conference (pp. 125\u2013130)"},{"key":"9929_CR37","doi-asserted-by":"publisher","unstructured":"Kuznetsova, A., Brockhoff, P. B., & Christensen, R. H. B. (2017). lmerTest package: Tests in linear mixed effects models. Journal of Statistical Software, 82(13), 1\u201326. https:\/\/doi.org\/10.18637\/jss.v082.i13","DOI":"10.18637\/jss.v082.i13"},{"key":"9929_CR38","doi-asserted-by":"crossref","unstructured":"Liberman, M. Y. (2019). Corpus phonetics. Annual Review of Linguistics, 5(1), 91\u2013107.","DOI":"10.1146\/annurev-linguistics-011516-033830"},{"key":"9929_CR39","doi-asserted-by":"crossref","unstructured":"Liu, H., MacWhinney, B., Fromm, D., & Lanzi, A. (2023). Automation of language sample analysis. Journal of Speech, Language, and Hearing Research, 66(7), 2421\u20132433.","DOI":"10.1044\/2023_JSLHR-22-00642"},{"key":"9929_CR40","doi-asserted-by":"crossref","unstructured":"Millasseau, J., Bruggeman, L., Yuen, I., & Demuth, K. (2021). Temporal cues to onset voicing contrasts in Australian English-speaking children. The Journal of the Acoustical Society of America, 149(1), 348\u2013356.","DOI":"10.1121\/10.0003060"},{"key":"9929_CR41","unstructured":"OAIC, A.G. (2026). What is personal information? Australian Government, Office of the Australian Information Commissioner. https:\/\/www.oaic.gov.au\/privacy\/your-privacy-rights\/your-personal-information\/what-is-personal-information"},{"key":"9929_CR42","doi-asserted-by":"crossref","unstructured":"Panayotov, V., Chen, G., Povey, D., & Khudanpur, S. (2015). Librispeech: an ASR corpus based on public domain audio books. In Proc. of IEEE international conference on acoustics, speech and signal processing (ICASSP) (pp. 5206\u20135210)","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"9929_CR43","doi-asserted-by":"crossref","unstructured":"Paulo, S., & Oliveira, L.C. (2004). Automatic phonetic alignment and its confidence measures. In Proc. of Advances in natural language processing  (pp. 36\u201344).","DOI":"10.1007\/978-3-540-30228-5_4"},{"key":"9929_CR44","doi-asserted-by":"crossref","unstructured":"P\u00e9rez-Espinosa, H., Mart\u00ednez-Miranda, J., Espinosa-Curiel, I., Rodr\u00edguez-Jacobo, J., Villase\u00f1or-Pineda, L., & Avila-George, H. (2020). IESC-child: an interactive emotional children\u2019s speech corpus. Computer Speech & Language, 59, 55\u201374.","DOI":"10.1016\/j.csl.2019.06.006"},{"key":"9929_CR45","doi-asserted-by":"crossref","unstructured":"Pitt, M. A., Johnson, K., Hume, E., Kiesling, S., & Raymond, W. (2005). The Buckeye corpus of conversational speech: Labeling conventions and a test of transcriber reliability. Speech Communication, 45(1), 89\u201395.","DOI":"10.1016\/j.specom.2004.09.001"},{"key":"9929_CR46","doi-asserted-by":"crossref","unstructured":"Potamianos, A., & Narayanan, S. (2003). Robust recognition of children\u2019s speech. IEEE Transactions on Speech and Audio Processing, 11(6), 603\u2013616.","DOI":"10.1109\/TSA.2003.818026"},{"key":"9929_CR47","unstructured":"Povey, D., Ghoshal, A., Boulianne, G., Burget, L., Glembek, O., Goel, N. & Vesely, K. (2011). The Kaldi speech recognition toolkit. IEEE 2011 workshop on automatic speech recognition and understanding."},{"key":"9929_CR48","doi-asserted-by":"crossref","unstructured":"Povey, D., Peddinti, V., Galvez, D., Ghahremani, P., Manohar, V., Na, X., & Khudanpur, S. (2016). Purely sequence-trained neural networks for ASR based on lattice-free MMI. In Proc. of INTERSPEECH (pp. 2751\u20132755).","DOI":"10.21437\/Interspeech.2016-595"},{"key":"9929_CR49","unstructured":"R Core Team (2025). Vienna, Austria. https:\/\/www.R-project.org\/"},{"key":"9929_CR50","doi-asserted-by":"crossref","unstructured":"Radha, K., Bansal, M., & Pachori, R. B. (2024). Speech and speaker recognition using raw waveform modeling for adult and children\u2019s speech: A comprehensive review. Engineering Applications of Artificial Intelligence, 131, Article 107661.","DOI":"10.1016\/j.engappai.2023.107661"},{"key":"9929_CR51","doi-asserted-by":"crossref","unstructured":"Rousseau, A., Del\u00e9glise, P., & Esteve, Y. (2014). Enhancing the TED-LIUM corpus with selected data for language modeling and more TED talks. LREC, 3935\u20133939.","DOI":"10.63317\/52t9tccogrod"},{"key":"9929_CR52","doi-asserted-by":"crossref","unstructured":"Rumberg, L., Gebauer, C., Ehlert, H., Wallbaum, M., Bornholt, L., Ostermann, J., & L\u00fcdtke, U. (2022). kidsTALC: A corpus of 3-to 11-year-old German children\u2019s connected natural speech. In Proc. of INTERSPEECH (pp. 5160\u20135164).","DOI":"10.21437\/Interspeech.2022-330"},{"key":"9929_CR53","doi-asserted-by":"crossref","unstructured":"Schultz, B. G., Tarigoppula, V. S. A., Noffs, G., Rojas, S., van der Walt, A., Grayden, D. B., & Vogel, A. P. (2021). Automatic speech recognition in neurodegenerative disease. International Journal of Speech Technology, 24(3), 771\u2013779.","DOI":"10.1007\/s10772-021-09836-w"},{"key":"9929_CR54","doi-asserted-by":"crossref","unstructured":"Shahin, M., Epps, J., & Ahmed, B. (2025). Phonological level wav2vec2-based mispronunciation detection and diagnosis method. Speech Communication, 173, Article 103249.","DOI":"10.1016\/j.specom.2025.103249"},{"key":"9929_CR55","doi-asserted-by":"crossref","unstructured":"Shahin, M., Lu, R., Epps, J., & Ahmed, B. (2020). UNSW system description for the shared task on automatic speech recognition for non-native children\u2019s speech. In Proc. of INTERSPEECH  (pp. 265\u2013268).","DOI":"10.21437\/Interspeech.2020-3111"},{"key":"9929_CR56","doi-asserted-by":"crossref","unstructured":"Shivakumar, P. G., & Georgiou, P. (2020). Transfer learning from adult to children for speech recognition: Evaluation, analysis and recommendations. Computer speech & language, 63, Article 101077.","DOI":"10.1016\/j.csl.2020.101077"},{"key":"9929_CR57","unstructured":"Shobaki, K., Hosom, J.-P., & Cole, R. (2000). The OGI kids\u2019 speech corpus and recognizers. Proc. of ICSLP (pp. 564\u2013567)."},{"key":"9929_CR58","doi-asserted-by":"crossref","unstructured":"Smithson, M., & Verkuilen, J. (2006). A better lemon squeezer? Maximum-likelihood regression with beta-distributed dependent variables. Psychological Methods, 11(1), 54\u201371.","DOI":"10.1037\/1082-989X.11.1.54"},{"key":"9929_CR59","doi-asserted-by":"crossref","unstructured":"Sobti, R., Guleria, K., & Kadyan, V. (2024). Comprehensive literature review on children automatic speech recognition system, acoustic linguistic mismatch approaches and challenges. Multimedia Tools and Applications, 83, 1\u201363.","DOI":"10.1007\/s11042-024-18753-4"},{"key":"9929_CR60","doi-asserted-by":"crossref","unstructured":"Sprugnoli, R., Moretti, G., Bentivogli, L., & Giuliani, D. (2017). Creating a ground truth multilingual dataset of news and talk show transcriptions through crowdsourcing. Language Resources and Evaluation, 51, 283\u2013317.","DOI":"10.1007\/s10579-016-9372-5"},{"key":"9929_CR61","doi-asserted-by":"crossref","unstructured":"Stevens, M., & Harrington, J. (2016). The phonetic origins of \/s\/-retraction: Acoustic and perceptual evidence from Australian English. Journal of Phonetics, 58, 118\u2013134.","DOI":"10.1016\/j.wocn.2016.08.003"},{"key":"9929_CR62","unstructured":"Szalay, T., Benders, T., Cox, F., Proctor, M. (2023). Pre-\/l\/ vowel change in Australian English pool-pull. In Proc. of the International Congress of Phonetic Sciences (pp. 2976\u20132980)."},{"key":"9929_CR63","doi-asserted-by":"crossref","unstructured":"Szalay, T., Nan, Z., Huang, R., Shahin, M., Sirojan, T., Ballard, K. J., & Ahmed, B. (2026). AusKidTalk: Developing transcription guidelines for continuous Australian English child speech. In Proc. of Language Resources and Evaluation Conference.  (pp. 5794\u20135804).","DOI":"10.63317\/4j6otnjq8c3n"},{"key":"9929_CR64","unstructured":"Szalay, T., Ratko, L., Shahin, M., Sirojan, T., Ballard, K.J., Cox, F., Ahmed, B. (2022). A semi-automatic workflow for orthographic transcription of a novel speech corpus: A case study of AusKidTalk. R.\u00a0Billington (Ed.), Proc. of Australasian International Conference on Speech Science and Technology (pp. 126\u2013130)."},{"key":"9929_CR65","doi-asserted-by":"crossref","unstructured":"Szalay, T., Shahin, M., Ahmed, B., & Ballard, K. J. (2022). Knowledge of accent differences can be used to predict speech recognition. In Proc. of INTERSPEECH (pp.1372\u20131376).","DOI":"10.21437\/Interspeech.2022-10162"},{"key":"9929_CR66","unstructured":"Szalay, T., Shahin, M., Ballard, K., Ahmed, B. (2022). Training forced aligners on (mis)matched data: the effect of dialect and age. R.\u00a0Billington (Ed.), Proc. of the  Australasian International Conference on Speech Science and Technology (pp. 36\u201340) ."},{"key":"9929_CR67","doi-asserted-by":"crossref","unstructured":"Szalay, T., Shahin, M., Sirojan, T., Nan, Z., Huang, R., Ballard, K.J., Ahmed, B. (2025). AusKidTalk: The benefits of out-of-domain automatic speech processing tools in corpus building. In Proc. of INTERSPEECH (pp. 4268-4272).","DOI":"10.21437\/Interspeech.2025-539"},{"key":"9929_CR68","doi-asserted-by":"crossref","unstructured":"Virkkunen, A., Rouhe, A., Phan, N., & Kurimo, M. (2023). Finnish parliament ASR corpus: Analysis, benchmarks and statistics. Language Resources and Evaluation, 57(4), 1645\u20131670.","DOI":"10.1007\/s10579-023-09650-7"},{"key":"9929_CR69","unstructured":"Wagner, M., Tran, D., Togneri, R., Rose, P., Powers, D., Onslow, M. & Ambikairajah, E. (2011). The big Australian speech corpus (the big ASC). Proc. of the Australasian International Conference on Speech Science and Technology (pp. 166\u2013170)."},{"key":"9929_CR70","unstructured":"Ward, W., Cole, R., & Pradhan, S. (2019). My science tutor and the MyST corpus. Carbondale: Boulder Learning Inc."},{"key":"9929_CR71","doi-asserted-by":"crossref","unstructured":"Wassink, A. B., Gansen, C., & Bartholomew, I. (2022). Uneven success: automatic speech recognition and ethnicity-related dialects. Speech Communication, 140, 50\u201370.","DOI":"10.1016\/j.specom.2022.03.009"},{"key":"9929_CR72","doi-asserted-by":"crossref","unstructured":"Werner-Seidler, A., Maston, K., Calear, A. L., Batterham, P. J., Larsen, M. E., Torok, M., et al. (2023). The future proofing study: Design, methods and baseline characteristics of a prospective cohort study of the mental health of Australian adolescents. International Journal of Methods in Psychiatric Research, 32(3), Article e1954.","DOI":"10.1002\/mpr.1954"},{"key":"9929_CR73","doi-asserted-by":"crossref","unstructured":"Yadav, I. C., Kumar, A., Shahnawazuddin, S., & Pradhan, G. (2018). Non-uniform spectral smoothing for robust children\u2019s speech recognition. In Proc. of INTERSPEECH  (pp. 1601\u20131605).","DOI":"10.21437\/Interspeech.2018-1828"}],"container-title":["Language Resources and Evaluation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10579-026-09929-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10579-026-09929-5","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10579-026-09929-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T03:52:14Z","timestamp":1783309934000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10579-026-09929-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,6]]},"references-count":73,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,9]]}},"alternative-id":["9929"],"URL":"https:\/\/doi.org\/10.1007\/s10579-026-09929-5","relation":{},"ISSN":["1574-020X","1574-0218"],"issn-type":[{"value":"1574-020X","type":"print"},{"value":"1574-0218","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,6]]},"assertion":[{"value":"3 June 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 June 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 July 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Co-authors Joanne Arciuli, Elise Baker, Felicity Cox and Kirrie Ballard own the intellectual property of the Speech Assessment for Australia Children (STAC) that is mentioned in the manuscript. They have no financial interests. The other authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"53"}}