{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:47:00Z","timestamp":1760237220667,"version":"build-2065373602"},"reference-count":60,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2020,3,13]],"date-time":"2020-03-13T00:00:00Z","timestamp":1584057600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Steady-state vowels are vowels that are uttered with a momentarily fixed vocal tract configuration and with steady vibration of the vocal folds. In this steady-state, the vowel waveform appears as a quasi-periodic string of elementary units called pitch periods. Humans perceive this quasi-periodic regularity as a definite pitch. Likewise, so-called pitch-synchronous methods exploit this regularity by using the duration of the pitch periods as a natural time scale for their analysis. In this work, we present a simple pitch-synchronous method using a Bayesian approach for estimating formants that slightly generalizes the basic approach of modeling the pitch periods as a superposition of decaying sinusoids, one for each vowel formant, by explicitly taking into account the additional low-frequency content in the waveform which arises not from formants but rather from the glottal pulse. We model this low-frequency content in the time domain as a polynomial trend function that is added to the decaying sinusoids. The problem then reduces to a rather familiar one in macroeconomics: estimate the cycles (our decaying sinusoids) independently from the trend (our polynomial trend function); in other words, detrend the waveform of steady-state waveforms. We show how to do this efficiently.<\/jats:p>","DOI":"10.3390\/e22030331","type":"journal-article","created":{"date-parts":[[2020,3,17]],"date-time":"2020-03-17T09:27:41Z","timestamp":1584437261000},"page":"331","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Detrending the Waveforms of Steady-State Vowels"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1935-1770","authenticated-orcid":false,"given":"Marnix","family":"Van Soom","sequence":"first","affiliation":[{"name":"Artificial Intelligence Laboratory, Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussels, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bart","family":"de Boer","sequence":"additional","affiliation":[{"name":"Artificial Intelligence Laboratory, Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussels, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,3,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Van Soom, M., and de Boer, B. (2019). A New Approach to the Formant Measuring Problem. Proceedings, 33.","DOI":"10.3390\/proceedings2019033029"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Fulop, S.A. (2011). Speech Spectrum Analysis, Springer. OCLC: 746243279; Signals and Communication Technology.","DOI":"10.1007\/978-3-642-17478-0"},{"key":"ref_3","unstructured":"Fant, G. (1960). Acoustic Theory of Speech Production, Mouton."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Stevens, K.N. (2000). Acoustic Phonetics, MIT Press.","DOI":"10.7551\/mitpress\/1072.001.0001"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Rabiner, L.R., and Schafer, R.W. (2007). Introduction to Digital Speech Processing, Foundations and Trends in Signal Processing.","DOI":"10.1561\/9781601980717"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Rose, P. (2002). Forensic Speaker Identification, CRC Press.","DOI":"10.1201\/9780203166369"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"894","DOI":"10.1016\/j.sleep.2007.07.010","article-title":"Could Formant Frequencies of Snore Signals Be an Alternative Means for the Diagnosis of Obstructive Sleep Apnea?","volume":"9","author":"Ng","year":"2008","journal-title":"Sleep Med."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Singh, R., Raj, B., and Gencaga, D. (June, January 30). Forensic Anthropometry from Voice: An Articulatory-Phonetic Approach. Proceedings of the 2016 39th International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO), Opatija, Croatia.","DOI":"10.1109\/MIPRO.2016.7522354"},{"key":"ref_9","unstructured":"Bretthorst, G.L. (2003). Probability Theory: The Logic of Science, Cambridge University Press."},{"key":"ref_10","unstructured":"Bonastre, J.F., Kahn, J., Rossato, S., and Ajili, M. (2020, March 12). Forensic Speaker Recognition: Mirages and Reality. Available online: https:\/\/www.oapen.org\/download?type=document&docid=1002748#page=257."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Hughes, N., and Karabiyik, U. (2020). Towards reliable digital forensics investigations through measurement science. WIREs Forensic Sci., e1367.","DOI":"10.1002\/wfs2.1367"},{"key":"ref_12","unstructured":"De Witte, W. (2017). A Forensic Speaker Identification Study: An Auditory-Acoustic Analysis of Phonetic Features and an Exploration of the \u201cTelephone Effect\u201d. [Ph.D. Thesis, Universitat Aut\u00f2noma de Barcelona]."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"74","DOI":"10.1016\/j.jcomdis.2018.05.004","article-title":"Static Measurements of Vowel Formant Frequencies and Bandwidths: A Review","volume":"74","author":"Kent","year":"2018","journal-title":"J. Commun. Disord."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"944","DOI":"10.1121\/1.4906840","article-title":"Statistical Properties of Linear Prediction Analysis Underlying the Challenge of Formant Bandwidth Estimation","volume":"137","author":"Mehta","year":"2015","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_15","unstructured":"Harrison, P. (2013). Making Accurate Formant Measurements: An Empirical Investigation of the Influence of the Measurement Tool, Analysis Settings and Speaker on Formant Measurements. [Ph.D. Thesis, University of York]."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Maurer, D. (2016). Acoustics of the Vowel, Peter Lang.","DOI":"10.3726\/978-3-0343-2391-8"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"713","DOI":"10.1121\/1.4940665","article-title":"Comparing Measurement Errors for Formants in Synthetic and Natural Vowels","volume":"139","author":"Shadle","year":"2016","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"38","DOI":"10.3390\/axioms1010038","article-title":"Foundations of Inference","volume":"1","author":"Knuth","year":"2012","journal-title":"Axioms"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1264","DOI":"10.1121\/1.1918682","article-title":"Pitch-Synchronous Time-Domain Estimation of Formant Frequencies and Bandwidths","volume":"35","author":"Pinson","year":"1963","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Mohammad-Djafari, A., and Demoment, G. (1993). Speech Processing Using Bayesian Inference. Maximum Entropy and Bayesian Methods: Paris, France, 1992, Springer. Fundamental Theories of Physics.","DOI":"10.1007\/978-94-017-2217-9"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"598","DOI":"10.1109\/TASL.2012.2229979","article-title":"Default Bayesian Estimation of the Fundamental Frequency","volume":"21","author":"Nielsen","year":"2013","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1044\/jshr.0901.68","article-title":"The Elements of an Acoustic Phonetic Theory","volume":"9","author":"Peterson","year":"1966","journal-title":"J. Speech Hear. Res."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1186\/1475-925X-6-23","article-title":"Exploiting Nonlinear Recurrence and Fractal Scaling Properties for Voice Disorder Detection","volume":"6","author":"Little","year":"2007","journal-title":"Biomed. Eng. Online"},{"key":"ref_24","unstructured":"Titze, I.R. (1995). Workshop on Acoustic Voice Analysis: Summary Statement, National Center for Voice and Speech."},{"key":"ref_25","unstructured":"While our (and others\u2019 [27]) everyday experience of looking at speech waveforms confirms this, we would be very interested in a formal study on this subject."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"582","DOI":"10.1007\/BF01789729","article-title":"Phonophotographische Untersuchungen","volume":"45","author":"Hermann","year":"1889","journal-title":"Pfl\u00fcg. Arch."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Chen, C.J. (2016). Elements of Human Voice, World Scientific.","DOI":"10.1142\/9891"},{"key":"ref_28","unstructured":"Scripture, E.W. (1904). The Elements of Experimental Phonetics, C. Scribner\u2019s Sons."},{"key":"ref_29","unstructured":"Kominek, J., and Black, A.W. (2020, March 12). The CMU Arctic Speech Databases. Available online: http:\/\/festvox.org\/cmu_arctic\/cmu_arctic_report.pdf."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Ladefoged, P. (1996). Elements of Acoustic Phonetics, University of Chicago Press.","DOI":"10.7208\/chicago\/9780226191010.001.0001"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"994","DOI":"10.1109\/TASL.2011.2170835","article-title":"Detection of Glottal Closure Instants From Speech Signals: A Quantitative Review","volume":"20","author":"Drugman","year":"2012","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_32","first-page":"40","article-title":"The LF-model revisited. Transformations and frequency domain analysis","volume":"2","author":"Fant","year":"1995","journal-title":"Speech Trans. Lab. Q. Rep. R. Inst. Tech. Stockh."},{"key":"ref_33","unstructured":"Thus, when a vowel is perceived with a clear and constant pitch, it is reasonable to assume that the vowel attained steady-state at some point (though perceptual effects forbid a one-to-one correspondence)."},{"key":"ref_34","first-page":"1","article-title":"Formant Bandwidth Data","volume":"3","author":"Fant","year":"1962","journal-title":"STL-QPSR"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1044\/jshr.0104.309","article-title":"Estimation of Formant Band Widths from Measurements of Transient Response of the Vocal Tract","volume":"1","author":"House","year":"1958","journal-title":"J. Speech Hear. Res."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Sanchez, J. (1989). Application of Classical, Bayesian and Maximum Entropy Spectrum Analysis to Nonstationary Time Series Data. Maximum Entropy and Bayesian Methods, Springer.","DOI":"10.1007\/978-94-015-7860-8_31"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"833","DOI":"10.1214\/06-BA127","article-title":"Nested Sampling for General Bayesian Computation","volume":"1","author":"Skilling","year":"2006","journal-title":"Bayesian Anal."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"\u00d3Ruanaidh, J.J.K., and Fitzgerald, W.J. (1996). Numerical Bayesian Methods Applied to Signal Processing, Springer. Statistics and Computing.","DOI":"10.1007\/978-1-4612-0717-7"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Bretthorst, G.L. (1988). Bayesian Spectrum Analysis and Parameter Estimation, Springer Science & Business Media.","DOI":"10.1007\/978-1-4684-9399-3"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Jaynes, E.T. (1987). Bayesian Spectrum and Chirp Analysis. Maximum-Entropy and Bayesian Spectral Analysis and Estimation Problems, Reidel.","DOI":"10.1007\/978-94-009-3961-5_1"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Sivia, D., and Skilling, J. (2006). Data Analysis: A Bayesian Tutorial, OUP.","DOI":"10.1093\/oso\/9780198568315.001.0001"},{"key":"ref_42","unstructured":"These are improved versions of the ones used in the conference paper."},{"key":"ref_43","unstructured":"This makes no difference in the posterior distribution for \u03b8 because the Legendre functions are linear combinations of the polynomials."},{"key":"ref_44","unstructured":"In our experiments, this regularizer prevented a situation that can best be described as polynomial and sinusoidal basis functions with huge amplitudes conspiring together into creating beats that, added together, fitted the data quite well but would yield nonphysical values for the \u03b8."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"175","DOI":"10.1121\/1.1906875","article-title":"Control Methods Used in a Study of the Vowels","volume":"24","author":"Peterson","year":"1952","journal-title":"J. Acoust. Soc. Am."},{"key":"ref_46","unstructured":"Another approach which avoids setting these prior ranges explicitly uses the perceived reliability of the LPC estimates to assign an expected relative accuracy \u03c1 (e.g., 10%) to the LPC estimates \u03b8^LPC. It is then possible to assign lognormal prior pdfs for the \u03b8 which are parametrized by setting the mean and standard deviation of the underlying normal distributions to log\u03b8^LPC and \u03c1, respectively [41]. This works well as long as \u03c1 \u2264 0.40, regardless of the value of \u03b8^LPC. We use the same technique in Section 4.1 to simulate errors in pitch period segmentation."},{"key":"ref_47","first-page":"28","article-title":"The Voice Source-Acoustic Modeling","volume":"4","author":"Fant","year":"1982","journal-title":"STL-QPSR"},{"key":"ref_48","unstructured":"Rabiner, L.R., Juang, B.H., and Rutledge, J.C. (1993). Fundamentals of Speech Recognition, PTR Prentice Hall Englewood Cliffs."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1016\/0167-6393(83)90064-X","article-title":"On the Time Domain Properties of the Two-Pole Model of the Glottal Waveform and Implications for LPC","volume":"2","author":"Deller","year":"1983","journal-title":"Speech Commun."},{"key":"ref_50","unstructured":"Petersen, K., and Pedersen, M. (2008). The Matrix Cookbook, Technical University of Denmark."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"R1","DOI":"10.1088\/0266-5611\/19\/2\/201","article-title":"Separable nonlinear least squares: The variable projection method and its applications","volume":"19","author":"Golub","year":"2003","journal-title":"Inverse Prob."},{"key":"ref_52","first-page":"341","article-title":"Praat, a system for doing phonetics by computer","volume":"5","author":"Boersma","year":"2001","journal-title":"Glot. Int."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.wocn.2018.07.001","article-title":"Introducing Parselmouth: A Python interface to Praat","volume":"71","author":"Jadoul","year":"2018","journal-title":"J. Phonetics"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Speagle, J.S. (2019). dynesty: A dynamic nested sampling package for estimating Bayesian posteriors and evidences. arXiv.","DOI":"10.1093\/mnras\/staa278"},{"key":"ref_55","unstructured":"Vall\u00e9e, N. (1994). Syst\u00e8mes Vocaliques: De La Typologie Aux Pr\u00e9dictions. [Ph.D. Thesis, l\u2019Universit\u00e9 Stendhal]."},{"key":"ref_56","first-page":"1","article-title":"A Four-Parameter Model of Glottal Flow","volume":"4","author":"Fant","year":"1985","journal-title":"STL-QPSR"},{"key":"ref_57","unstructured":"Praat\u2019s default formant estimation algorithm preprocesses the speech data by applying a +6 dB\/octave pre-emphasis filter to boost the amplitudes of the higher formants. The goal of this common technique [2] is to facilitate the measurement of the higher formants\u2019 bandwidths and frequencies. We have not applied pre-emphasis in our experiments, so it remains to be seen whether the third formant could have been picked up in this case. On that note, it might be of interest that this plus 6 dB\/oct pre-emphasis filter can be expressed in our Bayesian approach as prior covariance matrices {\u03a3i} for the noise vectors {ei} in Equation (5) specifying approximately a minus 6 dB\/oct slope for the prior noise spectral density. To see this, write the pre-emphasis operation on the data and model functions as: di \u2192 Eidi, Gi\u2192EiGi, (i = 1 \u22ef n), where Ei is a real and invertible Ni \u00d7 Ni matrix representing the pre-emphasis filter (for example, Praat by default uses yt \u2248 xt \u2212 0.98xt\u22121 which corresponds to Ei having ones on the principal diagonal and \u22120.98 on the subdiagonal). Then, the likelihood function in Equation (8) is proportional to exp{\u2212(1\/2) \u2211i=1n(di \u2212 Gibi)T\u2211i\u22121 (di \u2212 Gibi)} where the \u03a3i \u221d (EiT\n\t\t  Ei)\u22121 are positive definite covariance matrices specifying the prior for the spectral density of the noise, which turns out to be approximately \u22126 dB\/oct. Thus, pre-emphasis preprocessing can be interpreted as a more informative noise prior."},{"key":"ref_58","unstructured":"Titze, I.R. (2000). Principles of Voice Production (Second Printing), National Center for Voice and Speech."},{"key":"ref_59","unstructured":"Sundberg, J. (1978). Synthesis of singing. Swed. J. Musicol., 107\u2013112."},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Schroeder, M.R. (1999). Computer Speech: Recognition, Compression, Synthesis, Springer-Verlag.","DOI":"10.1007\/978-3-662-03861-1"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/22\/3\/331\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:06:48Z","timestamp":1760173608000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/22\/3\/331"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,13]]},"references-count":60,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2020,3]]}},"alternative-id":["e22030331"],"URL":"https:\/\/doi.org\/10.3390\/e22030331","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2020,3,13]]}}}