{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T18:45:43Z","timestamp":1776883543224,"version":"3.51.2"},"reference-count":56,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,3,27]],"date-time":"2024-03-27T00:00:00Z","timestamp":1711497600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,3,27]],"date-time":"2024-03-27T00:00:00Z","timestamp":1711497600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003816","name":"Huawei Technologies","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100003816","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Audio augmented reality (AAR), a prominent topic in the field of audio, requires understanding the listening environment of the user for rendering an authentic virtual auditory object. Reverberation time (<jats:inline-formula><jats:alternatives><jats:tex-math>$$RT_{60}$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mrow>\n                    <mml:mi>R<\/mml:mi>\n                    <mml:msub>\n                      <mml:mi>T<\/mml:mi>\n                      <mml:mn>60<\/mml:mn>\n                    <\/mml:msub>\n                  <\/mml:mrow>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>) is a predominant metric for the characterization of room acoustics and numerous approaches have been proposed to estimate it blindly from a reverberant speech signal. However, a single <jats:inline-formula><jats:alternatives><jats:tex-math>$$RT_{60}$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mrow>\n                    <mml:mi>R<\/mml:mi>\n                    <mml:msub>\n                      <mml:mi>T<\/mml:mi>\n                      <mml:mn>60<\/mml:mn>\n                    <\/mml:msub>\n                  <\/mml:mrow>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula> value may not be sufficient to correctly describe and render the acoustics of a room. This contribution presents a method for the estimation of multiple room acoustic parameters required to render close-to-accurate room acoustics in an unknown environment. It is shown how these parameters can be estimated blindly using an audio transformer that can be deployed on a mobile device. Furthermore, the paper also discusses the use of the estimated room acoustic parameters to find a similar room from a dataset of real BRIRs that can be further used for rendering the virtual audio source. Additionally, a novel binaural room impulse response (BRIR) augmentation technique to overcome the limitation of inadequate data is proposed. Finally, the proposed method is validated perceptually by means of a listening test.<\/jats:p>","DOI":"10.1186\/s13636-024-00338-6","type":"journal-article","created":{"date-parts":[[2024,3,27]],"date-time":"2024-03-27T16:02:07Z","timestamp":1711555327000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["An end-to-end approach for blindly rendering a virtual sound source in an audio augmented reality environment"],"prefix":"10.1186","volume":"2024","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-3981-2260","authenticated-orcid":false,"given":"Shivam","family":"Saini","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Isaac","family":"Engel","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"J\u00fcrgen","family":"Peissig","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,3,27]]},"reference":[{"issue":"3","key":"338_CR1","doi-asserted-by":"publisher","first-page":"63","DOI":"10.1109\/MSP.2021.3110108","volume":"39","author":"R Gupta","year":"2022","unstructured":"R. Gupta, J. He, R. Ranjan, W.S. Gan, F. Klein, C. Schneiderwind, A. Neidhardt, K. Brandenburg, V. V\u00e4lim\u00e4ki, Augmented\/mixed reality audio for hearables: Sensing, control, and rendering. IEEE Signal Proc. Mag. 39(3), 63\u201389 (2022). https:\/\/doi.org\/10.1109\/MSP.2021.3110108","journal-title":"IEEE Signal Proc. Mag."},{"issue":"6","key":"338_CR2","doi-asserted-by":"publisher","first-page":"4028","DOI":"10.1121\/1.2908264","volume":"123","author":"JE Summers","year":"2008","unstructured":"J.E. Summers, Auralization: Fundamentals of Acoustics, Modelling, Simulation, Algorithms, and Acoustic Virtual Reality. J. Acoust. Soc. Am. 123(6), 4028\u20134029 (2008). https:\/\/doi.org\/10.1121\/1.2908264","journal-title":"J. Acoust. Soc. Am."},{"key":"338_CR3","doi-asserted-by":"publisher","unstructured":"H. M\u00f8ller, Fundamentals of binaural technology. Appl. Acoust. 36(3), 171\u2013218 (1992). https:\/\/doi.org\/10.1016\/0003-682X(92)90046-U. https:\/\/www.sciencedirect.com\/science\/article\/pii\/0003682X9290046U","DOI":"10.1016\/0003-682X(92)90046-U"},{"key":"338_CR4","doi-asserted-by":"publisher","first-page":"111","DOI":"10.1121\/1.407089","volume":"94","author":"E Wenzel","year":"1993","unstructured":"E. Wenzel, M. Arruda, D. Kistler, F. Wightman, Localization using nonindividualized head-related transfer functions. J. Acoust. Soc. Am. 94, 111\u201323 (1993). https:\/\/doi.org\/10.1121\/1.407089","journal-title":"J. Acoust. Soc. Am."},{"key":"338_CR5","doi-asserted-by":"publisher","unstructured":"W. O. Brimijoin, A. W. Boyd, M. A. Akeroyd, The contribution of head movement to the externalization and internalization of sounds. PloS one.\u00a08(12), e83068 (2013). https:\/\/doi.org\/10.1371\/journal.pone.0083068","DOI":"10.1371\/journal.pone.0083068"},{"issue":"10","key":"338_CR6","first-page":"904","volume":"49","author":"DR Begault","year":"2001","unstructured":"D.R. Begault, E.M. Wenzel, M.R. Anderson, Direct comparison of the impact of head tracking, reverberation, and individualized head-related transfer functions on the spatial perception of a virtual speech source. J. Audio Eng. Soc. 49(10), 904\u2013916 (2001)","journal-title":"J. Audio Eng. Soc."},{"key":"338_CR7","doi-asserted-by":"crossref","unstructured":"S.\u00a0Werner, F.\u00a0Klein, T.\u00a0Mayenfels, K.\u00a0Brandenburg, in 2016 IEEE Eighth International Conference on Quality of Multimedia Experience (QoMEX), A summary on acoustic room divergence and its effect on externalization of auditory events (IEEE, 2016)","DOI":"10.1109\/QoMEX.2016.7498973"},{"key":"338_CR8","doi-asserted-by":"publisher","unstructured":"A.\u00a0Neidhardt, C.\u00a0Schneiderwind, F.\u00a0Klein, Perceptual matching of room acoustics for auditory augmented reality in small rooms - literature review and theoretical framework. Trends Hear. 26 (2022). https:\/\/doi.org\/10.1177\/23312165221092919","DOI":"10.1177\/23312165221092919"},{"issue":"4","key":"338_CR9","first-page":"219","volume":"49","author":"TJ Cox","year":"2001","unstructured":"T.J. Cox, F. Li, P. Darlington, Extracting room reverberation time from speech using artificial neural networks. J. Audio Eng. Soc. 49(4), 219\u2013230 (2001)","journal-title":"J. Audio Eng. Soc."},{"key":"338_CR10","unstructured":"H.\u00a0L\u00f6llmann, E.\u00a0Yilmaz, M.\u00a0Jeub, P.\u00a0Vary, in 2010 IEEE Proceedings of international workshop on acoustic echo and noise control (IWAENC), An improved algorithm for blind reverberation time estimation (IEEE, 2010)"},{"key":"338_CR11","unstructured":"L.\u00a0Treybig, S.\u00a0Saini, S.\u00a0Werner, U.\u00a0Sloma, J.\u00a0Peissig, in Audio Engineering Society Conference: AES 2022 International Audio for Virtual and Augmented Reality Conference, Room acoustic analysis and brir matching based on room acoustic measurements (Audio\u00a0Engineering Society, 2022)"},{"key":"338_CR12","doi-asserted-by":"crossref","unstructured":"J.\u00a0Eaton, N.D. Gaubitch, A.H. Moore, P.A. Naylor, in 2015 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), The ace challenge \u2014 corpus description and performance evaluation (IEEE, 2015)","DOI":"10.1109\/WASPAA.2015.7336912"},{"key":"338_CR13","doi-asserted-by":"publisher","unstructured":"S. Saini, J. Peissig, in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), Blind room acoustic parameters estimation using mobile audio transformer (2023). https:\/\/doi.org\/10.1109\/WASPAA58266.2023.10248186","DOI":"10.1109\/WASPAA58266.2023.10248186"},{"key":"338_CR14","doi-asserted-by":"publisher","unstructured":"M.\u00a0Lee, J.H. Chang, in 2016 IEEE International Conference on Network Infrastructure and Digital Content (IC-NIDC), Blind estimation of reverberation time using deep neural network. https:\/\/doi.org\/10.1109\/ICNIDC.2016.7974586","DOI":"10.1109\/ICNIDC.2016.7974586"},{"key":"338_CR15","doi-asserted-by":"publisher","unstructured":"H.\u00a0Gamper, I.J. Tashev, in 2018 16th International Workshop on Acoustic Signal Enhancement (IWAENC), Blind reverberation time estimation using a convolutional neural network. pp. 136\u2013140. https:\/\/doi.org\/10.1109\/IWAENC.2018.8521241","DOI":"10.1109\/IWAENC.2018.8521241"},{"issue":"2","key":"338_CR16","doi-asserted-by":"publisher","first-page":"255","DOI":"10.1109\/TASLP.2018.2877894","volume":"27","author":"F Xiong","year":"2019","unstructured":"F. Xiong, S. Goetze, B. Kollmeier, B.T. Meyer, Joint estimation of reverberation time and early-to-late reverberation ratio from single-channel speech signals. IEEE\/ACM Trans. Audio Speech Lang. Process. 27(2), 255\u2013267 (2019). https:\/\/doi.org\/10.1109\/TASLP.2018.2877894","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"338_CR17","doi-asserted-by":"publisher","unstructured":"D.\u00a0Looney, N.D. Gaubitch, in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Joint estimation of acoustic parameters from single-microphone speech observations. https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9054532","DOI":"10.1109\/ICASSP40776.2020.9054532"},{"key":"338_CR18","doi-asserted-by":"publisher","unstructured":"N.J. Bryan, in 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Impulse response data augmentation and deep neural networks for blind room acoustic parameter estimation. https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9052970","DOI":"10.1109\/ICASSP40776.2020.9052970"},{"key":"338_CR19","doi-asserted-by":"publisher","unstructured":"P.\u00a0G\u00f6tz, C.\u00a0Tuna, A.\u00a0Walther, E.A.P. Habets, in 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Blind reverberation time estimation in dynamic acoustic conditions. https:\/\/doi.org\/10.1109\/ICASSP43922.2022.9746457","DOI":"10.1109\/ICASSP43922.2022.9746457"},{"key":"338_CR20","doi-asserted-by":"publisher","unstructured":"S.\u00a0Deng, W.\u00a0Mack, E.A. Habets, in Proc. Interspeech 2020, Online Blind Reverberation Time Estimation Using CRNNs (2020), pp. 5061\u20135065. https:\/\/doi.org\/10.21437\/Interspeech.2020-2156","DOI":"10.21437\/Interspeech.2020-2156"},{"key":"338_CR21","doi-asserted-by":"publisher","unstructured":"C.\u00a0Ick, A.\u00a0Mehrabi, W.\u00a0Jin, in 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Blind acoustic room parameter estimation using phase features. https:\/\/doi.org\/10.1109\/ICASSP49357.2023.10094848","DOI":"10.1109\/ICASSP49357.2023.10094848"},{"key":"338_CR22","doi-asserted-by":"publisher","unstructured":"P. Srivastava, A. Deleforge, E. Vincent, in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA),\u00a0Blind room parameter estimation using multiple multichannel speech recordings (2021). https:\/\/doi.org\/10.1109\/WASPAA52581.2021.9632778","DOI":"10.1109\/WASPAA52581.2021.9632778"},{"key":"338_CR23","unstructured":"EN ISO 3382-2:2008 - Acoustics - Measurement of room acoustic parameters - Part 2: Reverberation time in ordinary rooms (ISO 3382-2:2008)"},{"issue":"10","key":"338_CR24","doi-asserted-by":"publisher","first-page":"1681","DOI":"10.1109\/TASLP.2016.2577502","volume":"24","author":"J Eaton","year":"2016","unstructured":"J. Eaton, N.D. Gaubitch, A.H. Moore, P.A. Naylor, Estimation of room acoustic parameters: The ace challenge. IEEE\/ACM Trans. Audio Speech Lang. Process. 24(10), 1681\u20131693 (2016). https:\/\/doi.org\/10.1109\/TASLP.2016.2577502","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"issue":"4","key":"338_CR25","doi-asserted-by":"publisher","first-page":"2251","DOI":"10.1121\/1.410097","volume":"96","author":"LG Marshall","year":"1994","unstructured":"L.G. Marshall, An acoustics measurement program for evaluating auditoriums based on the early\/late sound energy ratio. J. Acoust. Soc. Am. 96(4), 2251\u20132261 (1994)","journal-title":"J. Acoust. Soc. Am."},{"key":"338_CR26","doi-asserted-by":"publisher","unstructured":"H.\u00a0Gamper, in 2020 IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP), Blind c50 estimation from single-channel speech using a convolutional neural network. https:\/\/doi.org\/10.1109\/MMSP48831.2020.9287158","DOI":"10.1109\/MMSP48831.2020.9287158"},{"key":"338_CR27","unstructured":"P. Callens, M. Cernak, Joint blind room acoustic characterization from speech and music signals using convolutional recurrent neural networks.\u00a0(2020).\u00a0https:\/\/arxiv.org\/abs\/2010.11167\u00a0"},{"issue":"6","key":"338_CR28","doi-asserted-by":"publisher","first-page":"3532","DOI":"10.1121\/10.0019804","volume":"153","author":"P G\u00f6tz","year":"2023","unstructured":"P. G\u00f6tz, C. Tuna, A. Walther, E.A.P. Habets, Online reverberation time and clarity estimation in dynamic acoustic conditions. J. Acoust. Soc. Am. 153(6), 3532\u20133542 (2023). https:\/\/doi.org\/10.1121\/10.0019804","journal-title":"J. Acoust. Soc. Am."},{"key":"338_CR29","doi-asserted-by":"publisher","unstructured":"F.\u00a0Klein, A.\u00a0Neidhardt, M.\u00a0Seipel, Real-time estimation of reverberation time for selection of suitable binaural room impulse responses (2019). https:\/\/doi.org\/10.22032\/dbt.39968","DOI":"10.22032\/dbt.39968"},{"issue":"5","key":"338_CR30","doi-asserted-by":"publisher","first-page":"1991","DOI":"10.1109\/TVCG.2020.2973058","volume":"26","author":"Z Tang","year":"2020","unstructured":"Z. Tang, N.J. Bryan, D. Li, T.R. Langlois, D. Manocha, Scene-aware audio rendering via deep acoustic analysis. IEEE Trans. Vis. Comput. Graph. 26(5), 1991\u20132001 (2020). https:\/\/doi.org\/10.1109\/TVCG.2020.2973058","journal-title":"IEEE Trans. Vis. Comput. Graph."},{"key":"338_CR31","unstructured":"A.\u00a0Ratnarajah, S.\u00a0Ghosh, S.\u00a0Kumar, P.\u00a0Chiniya, D.\u00a0Manocha, Av-rir: Audio-visual room impulse response estimation (2023). arXiv preprint arXiv:2312.00834"},{"key":"338_CR32","doi-asserted-by":"crossref","unstructured":"C.J. Steinmetz, V.K. Ithapu, P.\u00a0Calamia, in 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), Filtered noise shaping for time domain room impulse response estimation from reverberant speech (IEEE, 2021)","DOI":"10.1109\/WASPAA52581.2021.9632680"},{"key":"338_CR33","unstructured":"A.\u00a0Ratnarajah, S.X. Zhang, Y.\u00a0Luo, D.\u00a0Yu, M3-audiodec: Multi-channel multi-speaker multi-spatial audio codec (2023). arXiv preprint arXiv:2309.07416"},{"key":"338_CR34","doi-asserted-by":"publisher","unstructured":"P. Li, Y. Song, I. McLoughlin, W. Guo, L. Dai, An Attention Pooling Based Representation Learning Method for Speech Emotion Recognition. Proc. Interspeech 2018, 3087\u20133091 (2018).\u00a0https:\/\/doi.org\/10.21437\/Interspeech.2018-1242","DOI":"10.21437\/Interspeech.2018-1242"},{"key":"338_CR35","doi-asserted-by":"publisher","unstructured":"Y.\u00a0Gong, Y.A. Chung, J.\u00a0Glass, in Proc. Interspeech 2021, AST: Audio Spectrogram Transformer (2021), p. 571\u2013575. https:\/\/doi.org\/10.21437\/Interspeech.2021-698","DOI":"10.21437\/Interspeech.2021-698"},{"key":"338_CR36","doi-asserted-by":"publisher","unstructured":"S.\u00a0Werner, F.\u00a0Klein, T.\u00a0Mayenfels, K.\u00a0Brandenburg, in 2016 Eighth International Conference on Quality of Multimedia Experience (QoMEX), A summary on acoustic room divergence and its effect on externalization of auditory events (2016), p. 1\u20136. https:\/\/doi.org\/10.1109\/QoMEX.2016.7498973","DOI":"10.1109\/QoMEX.2016.7498973"},{"key":"338_CR37","unstructured":"S.\u00a0Werner, G.\u00a0G\u00f6tz, F.\u00a0Klein, in Audio Engineering Society Convention 142, Influence of head tracking on the externalization of auditory events at divergence between synthesized and listening room using a binaural headphone system (Audio Engineering Society, 2017)"},{"key":"338_CR38","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-37762-4","volume-title":"The technology of binaural listening","author":"J Blauert","year":"2013","unstructured":"J. Blauert, The technology of binaural listening (Springer, Berlin, 2013)"},{"key":"338_CR39","volume-title":"in Audio Engineering Society Convention 129","author":"DT Murphy","year":"2010","unstructured":"D.T. Murphy, S. Shelley, in Audio Engineering Society Convention 129,\u00a0Openair: an interactive auralization web resource and database (Audio Engineering Society, 2010)"},{"issue":"4","key":"338_CR40","doi-asserted-by":"publisher","first-page":"863","DOI":"10.1109\/JSTSP.2019.2917582","volume":"13","author":"I Sz\u00f6ke","year":"2019","unstructured":"I. Sz\u00f6ke, M. Sk\u00e1cel, L. Mo\u0161ner, J. Paliesek, J. \u010cernock\u1ef3, Building and evaluation of a real room impulse response dataset. IEEE J. Sel. Top. Signal Process. 13(4), 863\u2013876 (2019)","journal-title":"IEEE J. Sel. Top. Signal Process."},{"issue":"8","key":"338_CR41","doi-asserted-by":"publisher","first-page":"1006","DOI":"10.1109\/LSP.2014.2379648","volume":"22","author":"GJ Mysore","year":"2014","unstructured":"G.J. Mysore, Can we automatically transform speech recorded on common consumer devices in real-world environments into professional production quality speech?\u2013a dataset, insights, and challenges. IEEE Signal Process. Lett. 22(8), 1006\u20131010 (2014)","journal-title":"IEEE Signal Process. Lett."},{"key":"338_CR42","doi-asserted-by":"publisher","unstructured":"C.\u00a0Hopkins, S.\u00a0Graetzer, G.\u00a0Seiffert. Aru speech corpus (University of Liverpool, 2019). https:\/\/doi.org\/10.17638\/datacat.liverpool.ac.uk\/681. https:\/\/datacat.liverpool.ac.uk\/681\/. Principal Investigator: Professor Carl Hopkins","DOI":"10.17638\/datacat.liverpool.ac.uk\/681"},{"key":"338_CR43","doi-asserted-by":"crossref","unstructured":"P. G\u00f6tz, C. Tuna, A. Walther, E.A. Habets, in\u00a0IEEE International Workshop on Acoustic Signal Enhancement (IWAENC),\u00a0Aid: Open-source anechoic interferer dataset (2022)","DOI":"10.1109\/IWAENC53105.2022.9914732"},{"key":"338_CR44","doi-asserted-by":"crossref","unstructured":"V.\u00a0Panayotov, G.\u00a0Chen, D.\u00a0Povey, S.\u00a0Khudanpur, in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), Librispeech: an asr corpus based on public domain audio books (IEEE, 2015)","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"338_CR45","doi-asserted-by":"publisher","first-page":"326","DOI":"10.1121\/1.2743161","volume":"122","author":"T Hidaka","year":"2007","unstructured":"T. Hidaka, Y. Yamada, T. Nakagawa, A new definition of boundary point between early reflections and late reverberation in room impulse responses. J. Acoust. Soc. Am. 122, 326\u201332 (2007). https:\/\/doi.org\/10.1121\/1.2743161","journal-title":"J. Acoust. Soc. Am."},{"key":"338_CR46","unstructured":"V.\u00a0Garcia-Gomez, J.J. Lopez, in Audio Engineering Society Convention 144, Binaural room impulse responses interpolation for multimedia real-time applications (Audio Engineering Society, 2018)"},{"key":"338_CR47","unstructured":"V.\u00a0Bruschi, S.\u00a0Nobili, A.\u00a0Terenzi, S.\u00a0Cecchi, in Audio Engineering Society Convention 152, An improved approach for binaural room impulse responses interpolation in real environments (Audio Engineering Society, 2022)"},{"key":"338_CR48","volume-title":"Partitioned convolution algorithms for real-time auralization","author":"F Wefers","year":"2015","unstructured":"F. Wefers, Partitioned convolution algorithms for real-time auralization, vol. 20 (Logos Verlag Berlin GmbH, Berlin, 2015)"},{"key":"338_CR49","doi-asserted-by":"crossref","unstructured":"T.d.M. Prego, A.A. de\u00a0Lima, R.\u00a0Zambrano-L\u00f3pez, S.L. Netto, in 2015 IEEE workshop on applications of signal processing to audio and acoustics (WASPAA), Blind estimators for reverberation time and direct-to-reverberant energy ratio using subband speech decomposition (IEEE, 2015)","DOI":"10.1109\/WASPAA.2015.7336954"},{"key":"338_CR50","unstructured":"J. Yamagishi, C. Veaux, K. MacDonald, CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92), University of Edinburgh. The Centre for Speech Technology Research (CSTR), (2019) Available: https:\/\/datashare.ed.ac.uk\/handle\/10283\/3443"},{"issue":"4","key":"338_CR51","first-page":"280","volume":"8","author":"HP Seraphim","year":"1958","unstructured":"H.P. Seraphim, Untersuchungen \u00fcber die unterschiedsschwelle exponentiellen abklingens von rauschbandimpulsen. Acta Acustica U. Acustica. 8(4), 280\u2013284 (1958)","journal-title":"Acta Acustica U. Acustica."},{"issue":"2","key":"338_CR52","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1016\/S0003-682X(98)00075-9","volume":"58","author":"JS Bradley","year":"1999","unstructured":"J.S. Bradley, R. Reich, S. Norcross, A just noticeable difference in c50 for speech. Appl. Acoust. 58(2), 99\u2013108 (1999)","journal-title":"Appl. Acoust."},{"key":"338_CR53","unstructured":"M. Blevins, A.T. Buck, Z. Peng, L. M. Wang, Quantifying the just noticeable difference of reverberation time with band-limited noise centered around 1000 Hz using a transformed up-down adaptive method (Proceedings of the International Symposium on Room Acoustics, Toronto, 2013)."},{"key":"338_CR54","unstructured":"International Telecom Union, Rec. ITU-R BS. 1534-1. Method for the subjective assessment of intermediate quality level of coding systems (2003)"},{"key":"338_CR55","unstructured":"International Telecom Union, Rec. ITU-R BS. 1534-3. Method for the subjective assessment of intermediate quality levels of coding systems (2015). https:\/\/www.itu.int\/rec\/R-REC-BS.1534"},{"key":"338_CR56","unstructured":"S.N. Wadekar, A.\u00a0Chaurasia, Mobilevitv3: mobile-friendly vision transformer with simple and effective fusion of local, global and input features (2022). arXiv preprint arXiv:2209.15159"}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-024-00338-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-024-00338-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-024-00338-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,3,27]],"date-time":"2024-03-27T16:06:25Z","timestamp":1711555585000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-024-00338-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,27]]},"references-count":56,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["338"],"URL":"https:\/\/doi.org\/10.1186\/s13636-024-00338-6","relation":{},"ISSN":["1687-4722"],"issn-type":[{"value":"1687-4722","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,27]]},"assertion":[{"value":"9 November 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 March 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 March 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"16"}}