{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,1]],"date-time":"2025-10-01T15:56:45Z","timestamp":1759334205199,"version":"build-2065373602"},"reference-count":20,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T00:00:00Z","timestamp":1759190400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T00:00:00Z","timestamp":1759190400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["SN COMPUT. SCI."],"DOI":"10.1007\/s42979-025-04383-6","type":"journal-article","created":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T12:13:59Z","timestamp":1759234439000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["PAZHVAK: a Word-Level Farsi Speech Corpus by University of Hormozgan"],"prefix":"10.1007","volume":"6","author":[{"given":"Mohammad Azim","family":"Saraji","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4006-2815","authenticated-orcid":false,"given":"Abdullah","family":"Khalili","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ahmad","family":"Hatam","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,9,30]]},"reference":[{"key":"4383_CR1","doi-asserted-by":"crossref","unstructured":"Prabhavalkar R, Hori T, Sainath TN, Schl\u00fcter R, Watanabe S. End-to-end speech recognition: a survey (2023). arXiv:2303.03329","DOI":"10.1109\/TASLP.2023.3328283"},{"key":"4383_CR2","unstructured":"Karmakar P, Teng SW, Lu G. Thank you for attention: a survey on attention-based artificial neural networks for automatic speech recognition 2021; arXiv:2102.07259"},{"key":"4383_CR3","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2024.102422","volume":"109","author":"H Kheddar","year":"2024","unstructured":"Kheddar H, Hemis M, Himeur Y. Automatic speech recognition using advanced deep learning approaches: a survey. Inf Fusion. 2024;109:102422. https:\/\/doi.org\/10.1016\/j.inffus.2024.102422.","journal-title":"Inf Fusion"},{"key":"4383_CR4","doi-asserted-by":"crossref","unstructured":"Bai Z, Zhang X-L. Speaker recognition based on deep learning: an overview 2021; arXiv:2012.00931","DOI":"10.1016\/j.neunet.2021.03.004"},{"key":"4383_CR5","unstructured":"Tan X, Qin T, Soong F, Liu T-Y. A survey on neural speech synthesis 2021; arXiv:2106.15561"},{"key":"4383_CR6","unstructured":"Chowdhury MJU, Hussan A. A review-based study on different text-to-Speech technologies 2023; arXiv:2312.11563"},{"key":"4383_CR7","unstructured":"Zhang C, Zhang C, Zheng S, Zhang M, Qamar M, Bae S-H, Kweon IS. A survey on audio diffusion models: text to speech synthesis and enhancement in generative AI. 2023. arXiv:2303.13336"},{"issue":"8","key":"4383_CR8","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter S, Schmidhuber J. Long short-term memory. Neural Comput. 1997;9(8):1735\u201380. https:\/\/doi.org\/10.1162\/neco.1997.9.8.1735.","journal-title":"Neural Comput"},{"key":"4383_CR9","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I. Attention is all you need. 2023; arXiv:1706.03762"},{"key":"4383_CR10","doi-asserted-by":"publisher","unstructured":"Panayotov V, Chen G, Povey D, Khudanpur S. Librispeech: an asr corpus based on public domain audio books. In: 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2015; 5206\u20135210. https:\/\/doi.org\/10.1109\/ICASSP.2015.7178964","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"4383_CR11","unstructured":"Warden P. Speech commands: a dataset for limited-vocabulary speech recognition. 2018; arXiv:1804.03209"},{"key":"4383_CR12","unstructured":"Ito K, Johnson L. The LJ speech dataset. https:\/\/keithito.com\/LJ-Speech-Dataset\/, 2017."},{"key":"4383_CR13","doi-asserted-by":"crossref","unstructured":"Zen H, Dang V, Clark R, Zhang Y, Weiss RJ, Jia Y, Chen Z, Wu Y. LibriTTS: a corpus derived from LibriSpeech for text-to-speech (2019). arXiv:1904.02882","DOI":"10.21437\/Interspeech.2019-2441"},{"key":"4383_CR14","doi-asserted-by":"publisher","unstructured":"MalekzadeH S, Gholizadeh MH, Razavi SN. Full persian vowel recognition with mfcc and ann on pcvc speech dataset. 2018. https:\/\/doi.org\/10.13140\/RG.2.2.12187.72486.","DOI":"10.13140\/RG.2.2.12187.72486"},{"key":"4383_CR15","unstructured":"Jia Y, Ramanovich MT, Wang Q, Zen H. CVSS corpus and massively multilingual speech-to-speech translation. 2022; arXiv:2201.03713"},{"key":"4383_CR16","doi-asserted-by":"crossref","unstructured":"Wang C, Wu A, Pino J. CoVoST 2 and massively multilingual speech-to-text translation. 2020; arXiv:2007.10310","DOI":"10.21437\/Interspeech.2021-2027"},{"key":"4383_CR17","doi-asserted-by":"crossref","unstructured":"Wang C, Pino J, Wu A, Gu J. CoVoST: a diverse multilingual speech-to-text translation corpus 2020; arXiv:2002.01320","DOI":"10.21437\/Interspeech.2021-2027"},{"key":"4383_CR18","unstructured":"Radford A, Kim JW, Xu T, Brockman G, McLeavey C, Sutskever I. Robust speech recognition via large-scale weak supervision 2022; arXiv:2212.04356"},{"key":"4383_CR19","unstructured":"Conneau A, Ma M, Khanuja S, Zhang Y, Axelrod V, Dalmia S, Riesa J, Rivera C, Bapna A. FLEURS: few-shot learning evaluation of universal representations of speech. 2022; arXiv:2205.12446"},{"key":"4383_CR20","unstructured":"Baevski A, Zhou H, Mohamed A, Auli M. wav2vec 2.0: a framework for self-supervised learning of speech representations 2020; arXiv:2006.11477"}],"container-title":["SN Computer Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s42979-025-04383-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s42979-025-04383-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s42979-025-04383-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T12:14:09Z","timestamp":1759234449000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s42979-025-04383-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,30]]},"references-count":20,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2025,10]]}},"alternative-id":["4383"],"URL":"https:\/\/doi.org\/10.1007\/s42979-025-04383-6","relation":{},"ISSN":["2661-8907"],"issn-type":[{"type":"electronic","value":"2661-8907"}],"subject":[],"published":{"date-parts":[[2025,9,30]]},"assertion":[{"value":"17 March 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 September 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 September 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not Applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Research Involving Human and\/or Animals"}},{"value":"This study was conducted in accordance with established ethical guidelines for research involving human participants. Prior to participation, all individuals were provided with clear and comprehensive information about the purpose, procedures, potential risks, and benefits of the study. Participants were given the opportunity to ask questions and were clearly notified that their involvement was entirely voluntary. In addition, all participants in this study provided informed consent prior to their involvement. They were clearly informed of their right to decline participation or withdraw from the study at any time without any negative consequences. To protect participants\u2019 privacy, all data were collected and stored anonymously. Personal identifiers were either not recorded or were removed during data processing to ensure confidentiality. Access to research data was restricted to authorized members of the research team only.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Informed Consent"}}],"article-number":"865"}}