{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T13:18:00Z","timestamp":1740143880432,"version":"3.37.3"},"reference-count":37,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,5,22]],"date-time":"2024-05-22T00:00:00Z","timestamp":1716336000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,5,22]],"date-time":"2024-05-22T00:00:00Z","timestamp":1716336000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["352015383, SFB 1330"],"award-info":[{"award-number":["352015383, SFB 1330"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100019559","name":"Carl von Ossietzky Universit\u00e4t Oldenburg","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100019559","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J AUDIO SPEECH MUSIC PROC."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Selective attention is a crucial ability of the auditory system. Computationally, following an auditory object can be illustrated as tracking its acoustic properties, e.g., pitch, timbre, or location in space. The difficulty is related to the fact that in a complex auditory scene, the information about the tracked object is not available in a clean form. The more cluttered the sound mixture, the more time and frequency regions where the object of interest is masked by other sound sources. How does the auditory system recognize and follow acoustic objects based on this fragmentary information? Numerous studies highlight the crucial role of top-down processing in this task. Having in mind both auditory modeling and signal processing applications, we investigated how computational methods with and without top-down processing deal with increasing sparsity of the auditory features in the task of estimating instantaneous voice states, defined as a combination of three parameters: fundamental frequency F0 and formant frequencies F1 and F2. We found that the benefit from top-down processing grows with increasing sparseness of the auditory data.<\/jats:p>","DOI":"10.1186\/s13636-024-00350-w","type":"journal-article","created":{"date-parts":[[2024,5,22]],"date-time":"2024-05-22T03:38:30Z","timestamp":1716349110000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Towards multidimensional attentive voice tracking\u2014estimating voice state from auditory glimpses with regression neural networks and Monte Carlo sampling"],"prefix":"10.1186","volume":"2024","author":[{"given":"Joanna","family":"Luberadzka","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hendrik","family":"Kayser","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"J\u00f6rg","family":"L\u00fccke","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Volker","family":"Hohmann","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,5,22]]},"reference":[{"issue":"3","key":"350_CR1","doi-asserted-by":"publisher","first-page":"1562","DOI":"10.1121\/1.2166600","volume":"119","author":"M Cooke","year":"2006","unstructured":"M. Cooke, A glimpsing model of speech perception in noise. J. Acoust. Soc. Am. 119(3), 1562\u20131573 (2006)","journal-title":"J. Acoust. Soc. Am."},{"key":"350_CR2","first-page":"73","volume-title":"in Physiology, psychoacoustics and cognition in normal and impaired hearing, Intelligibility for binaural speech with discarded low-snr speech components","author":"E Schoenmaker","year":"2016","unstructured":"E. Schoenmaker, S. van de Par. Intelligibility for binaural speech with discarded low-SNR speech components, in Physiology, psychoacoustics and cognition in normal and impaired hearing (Springer International Publishing, 2016), pp. 73\u201381"},{"key":"350_CR3","doi-asserted-by":"crossref","unstructured":"R.L. Gregory, Perceptions as hypotheses. Philos. Trans. R. Soc. Lond. B Biol. Sci. 290(1038), 181\u2013197 (1980)","DOI":"10.1098\/rstb.1980.0090"},{"issue":"2","key":"350_CR4","doi-asserted-by":"publisher","first-page":"712","DOI":"10.1121\/10.0009337","volume":"151","author":"J Luberadzka","year":"2022","unstructured":"J. Luberadzka, H. Kayser, V. Hohmann, Making sense of periodicity glimpses in a prediction-update-loop\u2013a computational model of attentive voice tracking. J Acoust. Soc. Am. 151(2), 712\u2013737 (2022)","journal-title":"J Acoust. Soc. Am."},{"issue":"17","key":"350_CR5","doi-asserted-by":"publisher","first-page":"2238","DOI":"10.1016\/j.cub.2015.07.043","volume":"25","author":"KJ Woods","year":"2015","unstructured":"K.J. Woods, J.H. McDermott, Attentive tracking of sound sources. Curr. Biol. 25(17), 2238\u20132246 (2015)","journal-title":"Curr. Biol."},{"issue":"5","key":"350_CR6","doi-asserted-by":"publisher","first-page":"2911","DOI":"10.1121\/1.4950699","volume":"139","author":"A Josupeit","year":"2016","unstructured":"A. Josupeit, N. Kop\u010do, V. Hohmann, Modeling of speech localization in a multi-talker mixture using periodicity and energy-based auditory features. J. Acoust. Soc. Am. 139(5), 2911\u20132923 (2016)","journal-title":"J. Acoust. Soc. Am."},{"issue":"1","key":"350_CR7","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1121\/1.4990375","volume":"142","author":"A Josupeit","year":"2017","unstructured":"A. Josupeit, V. Hohmann, Modeling speech localization, talker identification, and word recognition in a multi-talker setting. J. Acoust. Soc. Am. 142(1), 35\u201354 (2017)","journal-title":"J. Acoust. Soc. Am."},{"key":"350_CR8","volume-title":"Sparse periodicity-based auditory features explain human performance in a spatial multitalker auditory scene analysis task","author":"A Josupeit","year":"2018","unstructured":"A. Josupeit, E. Schoenmaker, S. van de Par, V. Hohmann, Sparse periodicity-based auditory features explain human performance in a spatial multitalker auditory scene analysis task. Eur. J. NeuroSci. (2018)"},{"key":"350_CR9","doi-asserted-by":"crossref","unstructured":"M.S Arulampalam, S. Maskell, N. Gordon, T. Clapp, A tutorial on particle filters for online nonlinear\/non-Gaussian Bayesian tracking. IEEE Transactions on signal processing. 50(2), 174\u201388 (2002)","DOI":"10.1109\/78.978374"},{"issue":"1","key":"350_CR10","doi-asserted-by":"publisher","first-page":"143","DOI":"10.3758\/s13423-016-1015-8","volume":"25","author":"D Van Ravenzwaaij","year":"2018","unstructured":"D. Van Ravenzwaaij, P. Cassey, S.D. Brown, A simple introduction to markov chain monte-carlo sampling. Psychon. Bull. Rev. 25(1), 143\u2013154 (2018)","journal-title":"Psychon. Bull. Rev."},{"issue":"6","key":"350_CR11","doi-asserted-by":"publisher","first-page":"568","DOI":"10.1109\/72.97934","volume":"2","author":"DF Specht","year":"1991","unstructured":"D.F. Specht et al., A general regression neural network. IEEE Trans. Neural Netw. 2(6), 568\u2013576 (1991)","journal-title":"IEEE Trans. Neural Netw."},{"key":"350_CR12","doi-asserted-by":"crossref","unstructured":"J. Luberadzka, H. Kayser, V. Hohmann. Estimating fundamental frequency and formants based on periodicity glimpses: A deep learning approach, in 2020 IEEE International Conference on Healthcare Informatics (ICHI), vol. 30 (IEEE, 2020), pp. 1\u20136","DOI":"10.1109\/ICHI48887.2020.9374386"},{"key":"350_CR13","unstructured":"V.\u00a0Hohmann. Method for extracting periodic signal components, and apparatus for this purpose (Google Patents, 2006). US Patent App. 11\/223125"},{"key":"350_CR14","unstructured":"J. Luberadzka, H. Kayser, V. Hohmann, Glimpsed periodicity features and recursive Bayesian estimation for modeling attentive voice tracking. Universit\u00e4tsbibliothek der RWTH Aachen; 2019."},{"issue":"1","key":"350_CR15","first-page":"1","volume":"182","author":"Z Chen","year":"2003","unstructured":"Z. Chen et al., Bayesian filtering: From kalman filters to particle filters, and beyond. Stat. 182(1), 1\u201369 (2003)","journal-title":"Stat."},{"key":"350_CR16","unstructured":"S.\u00a0Ruder, An overview of gradient descent optimization algorithms (2016). arXiv preprint arXiv:1609.04747"},{"issue":"3","key":"350_CR17","doi-asserted-by":"publisher","first-page":"971","DOI":"10.1121\/1.383940","volume":"67","author":"DH Klatt","year":"1980","unstructured":"D.H. Klatt, Software for a cascade\/parallel formant synthesizer. J. Acoust. Soc. Am. 67(3), 971\u2013995 (1980)","journal-title":"J. Acoust. Soc. Am."},{"issue":"6","key":"350_CR18","doi-asserted-by":"publisher","first-page":"3323","DOI":"10.1121\/1.1572146","volume":"113","author":"JG Bernstein","year":"2003","unstructured":"J.G. Bernstein, A.J. Oxenham, Pitch discrimination of diotic and dichotic tone complexes: harmonic resolvability or harmonic number? J. Acoust. Soc. Am. 113(6), 3323\u20133334 (2003)","journal-title":"J. Acoust. Soc. Am."},{"key":"350_CR19","unstructured":"S. Mittal, A. Lamb, A. Goyal, V. Voleti, M. Shanahan, G. Lajoie, M. Mozer, Y. Bengio. Learning to combine top-down and bottom-up signals in recurrent neural networks with attention over modules, in International Conference on Machine Learning, vol. 21 (PMLR, 2020), pp. 6972\u20136986"},{"issue":"3","key":"350_CR20","doi-asserted-by":"publisher","first-page":"479","DOI":"10.1016\/S0893-6080(96)00062-7","volume":"10","author":"D Husmeier","year":"1997","unstructured":"D. Husmeier, J.G. Taylor, Predicting conditional probability densities of stationary stochastic time series. Neural Netw. 10(3), 479\u2013497 (1997)","journal-title":"Neural Netw."},{"issue":"2194","key":"350_CR21","doi-asserted-by":"publisher","first-page":"20200209","DOI":"10.1098\/rsta.2020.0209","volume":"379","author":"B Lim","year":"2021","unstructured":"B. Lim, S. Zohren, Time-series forecasting with deep learning: a survey. Phil. Trans. R. Soc. A. 379(2194), 20200209 (2021)","journal-title":"Phil. Trans. R. Soc. A."},{"key":"350_CR22","unstructured":"M.F. Stollenga, J. Masci, F. Gomez, J. Schmidhuber, Deep networks with internal selective attention through feedback connections. Advances in neural information processing systems 27, (2014)"},{"key":"350_CR23","unstructured":"D.\u00a0Bahdanau, K.\u00a0Cho, Y.\u00a0Bengio, Neural machine translation by jointly learning to align and translate (2014). arXiv preprint arXiv:1409.0473"},{"key":"350_CR24","doi-asserted-by":"crossref","unstructured":"A. Rosenfeld, M. Biparva, J.K. Tsotsos. Priming neural networks, in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (2018), pp. 2011\u20132020","DOI":"10.1109\/CVPRW.2018.00270"},{"key":"350_CR25","unstructured":"D.J. Rezende, S. Mohamed, D. Wierstra. Stochastic backpropagation and approximate inference in deep generative models, in International conference on machine learning (PMLR, 2014),  pp. 1278\u20131286"},{"key":"350_CR26","unstructured":"Kingma DP, Welling M. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114. (2013)"},{"key":"350_CR27","unstructured":"S. Ramchandran, G. Tikhonov, K. Kujanp\u00e4\u00e4, M. Koskinen, H. L\u00e4hdesm\u00e4ki. Longitudinal variational autoencoder, in International Conference on Artificial Intelligence and Statistics (PMLR, 2021), pp. 3898\u20133906"},{"key":"350_CR28","unstructured":"V. Fortuin, D. Baranchuk, G. R\u00e4tsch, S. Mandt. Gp-vae: Deep probabilistic time series imputation, in International conference on artificial intelligence and statistics (PMLR, 2020), pp. 1651\u20131661"},{"key":"350_CR29","unstructured":"M.\u00a0Ashman, J.\u00a0So, W.\u00a0Tebbutt, V.\u00a0Fortuin, M.\u00a0Pearce, R.E. Turner, Sparse gaussian process variational autoencoders (2020). arXiv preprint arXiv:2010.10177"},{"key":"350_CR30","doi-asserted-by":"publisher","first-page":"107501","DOI":"10.1016\/j.patcog.2020.107501","volume":"107","author":"A Nazabal","year":"2020","unstructured":"A. Nazabal, P.M. Olmos, Z. Ghahramani, I. Valera, Handling incomplete heterogeneous data using VAEs. Pattern Recognit. 107, 107501 (2020)","journal-title":"Pattern Recognit."},{"issue":"6","key":"350_CR31","doi-asserted-by":"publisher","first-page":"4007","DOI":"10.1121\/1.2363929","volume":"120","author":"DS Brungart","year":"2006","unstructured":"D.S. Brungart, P.S. Chang, B.D. Simpson, D. Wang, Isolating the energetic component of speech-on-speech masking with ideal time-frequency segregation. J. Acoust. Soc. Am. 120(6), 4007\u20134018 (2006)","journal-title":"J. Acoust. Soc. Am."},{"key":"350_CR32","doi-asserted-by":"crossref","unstructured":"E.d. Boer, Pitch of inharmonic signals. Nat. 178(4532), 535\u2013536 (1956)","DOI":"10.1038\/178535a0"},{"issue":"9B","key":"350_CR33","doi-asserted-by":"publisher","first-page":"1418","DOI":"10.1121\/1.1918360","volume":"34","author":"JF Schouten","year":"1962","unstructured":"J.F. Schouten, R. Ritsma, B.L. Cardozo, Pitch of the residue. J. Acoust. Soc. Am. 34(9B), 1418\u20131424 (1962)","journal-title":"J. Acoust. Soc. Am."},{"key":"350_CR34","doi-asserted-by":"crossref","unstructured":"P.A. Cariani, B.\u00a0Delgutte, Neural correlates of the pitch of complex tones. ii. pitch shift, pitch ambiguity, phase invariance, pitch circularity, rate pitch, and the dominance region for pitch. J. Neurophys. 76(3), 1717\u20131734 (1996)","DOI":"10.1152\/jn.1996.76.3.1717"},{"key":"350_CR35","unstructured":"E Terhardt. On the role of ambiguity of perceived pitch in music, in Proc. 13th ICA Belgrade (1989), pp. 35\u201338"},{"issue":"2","key":"350_CR36","doi-asserted-by":"publisher","first-page":"680","DOI":"10.1121\/1.399772","volume":"88","author":"PF Assmann","year":"1990","unstructured":"P.F. Assmann, Q. Summerfield, Modeling the perception of concurrent vowels: vowels with different fundamental frequencies. J. Acoust. Soc. Am. 88(2), 680\u2013697 (1990)","journal-title":"J. Acoust. Soc. Am."},{"key":"350_CR37","doi-asserted-by":"crossref","unstructured":"M.R. Saddler, R. Gonzalez,  J.H. McDermott, Deep neural network models reveal interplay of peripheral coding and stimulus statistics in pitch perception. Nature communications. 12(1), 7278 (2021)","DOI":"10.1038\/s41467-021-27366-6"}],"container-title":["EURASIP Journal on Audio, Speech, and Music Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-024-00350-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s13636-024-00350-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s13636-024-00350-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,22]],"date-time":"2024-05-22T03:42:20Z","timestamp":1716349340000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmp-eurasipjournals.springeropen.com\/articles\/10.1186\/s13636-024-00350-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,22]]},"references-count":37,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["350"],"URL":"https:\/\/doi.org\/10.1186\/s13636-024-00350-w","relation":{},"ISSN":["1687-4722"],"issn-type":[{"type":"electronic","value":"1687-4722"}],"subject":[],"published":{"date-parts":[[2024,5,22]]},"assertion":[{"value":"22 November 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 May 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 May 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"27"}}