{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,31]],"date-time":"2026-01-31T04:09:53Z","timestamp":1769832593227,"version":"3.49.0"},"reference-count":23,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2023,6,1]],"date-time":"2023-06-01T00:00:00Z","timestamp":1685577600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,6,1]],"date-time":"2023-06-01T00:00:00Z","timestamp":1685577600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Tokyo University of Science"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Life Robotics"],"published-print":{"date-parts":[[2023,8]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>According to a survey on the cause of death among Japanese people, lifestyle-related diseases (such as malignant neoplasms, cardiovascular diseases, and pneumonia) account for 55.8% of all deaths. Three habits, namely, drinking, smoking, and sleeping, are considered the most important factors associated with lifestyle-related diseases, but it is difficult to measure these habits autonomously and regularly. Here, we propose a machine learning-based approach for detecting these lifestyle habits using voice data. We used classifiers and probabilistic linear discriminant analysis based on acoustic features, such as mel-frequency cepstrum coefficients (MFCCs) and jitter, extracted from a speech dataset we developed, and an X-vector from a pre-trained ECAPA-TDNN model. For training models, we used several classifiers implemented in MATLAB 2021b, such as support vector machines, K-nearest neighbors (KNN), and ensemble methods with some feature-projection options. Our results show that a cubic KNN method using acoustic features performs well on the sleep habit classification, while X-vector-based models perform well on smoking and drinking habit classifications. These results suggest that X-vectors may help estimate factors directly affecting the vocal cords and vocal tracts of the users (e.g., due to smoking and drinking), while acoustic features may help classify chronotypes, which might be informative with respect to the individuals\u2019 vocal cord and vocal tract ultrastructure.<\/jats:p>","DOI":"10.1007\/s10015-023-00870-2","type":"journal-article","created":{"date-parts":[[2023,6,1]],"date-time":"2023-06-01T15:02:02Z","timestamp":1685631722000},"page":"520-529","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Estimation of habit-related information from male voice data using machine learning-based methods"],"prefix":"10.1007","volume":"28","author":[{"given":"Takaya","family":"Yokoo","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ryo","family":"Hatano","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hiroyuki","family":"Nishiyama","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,6,1]]},"reference":[{"key":"870_CR1","unstructured":"Alcohol Health and Medical Association: Alcohol blood levels and drunkenness. http:\/\/www.arukenkyo.or.jp\/health\/base\/index.html Accessed 12 Nov 2021"},{"key":"870_CR2","doi-asserted-by":"crossref","unstructured":"Chung JS, Nagrani A, Zisserman A (2018) Voxceleb2: Deep speaker recognition. arXiv preprint arXiv:1806.05622","DOI":"10.21437\/Interspeech.2018-1929"},{"key":"870_CR3","doi-asserted-by":"crossref","unstructured":"Desplanques B, Thienpondt J, Demuynck K (2020) ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification. arXiv preprint arXiv:2005.07143","DOI":"10.21437\/Interspeech.2020-2650"},{"key":"870_CR4","doi-asserted-by":"crossref","unstructured":"Doukhan D, Carrive J, Vallet F, Larcher A, Meignier S (2018) An open-source speaker gender detection framework for monitoring gender equality. In: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5214\u20135218. 10.1109\/ICASSP.2018.8461471","DOI":"10.1109\/ICASSP.2018.8461471"},{"issue":"1","key":"870_CR5","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40345-021-00243-3","volume":"9","author":"M Faurholt-Jepsen","year":"2021","unstructured":"Faurholt-Jepsen M, Rohani DA, Busk J, Vinberg M, Bardram JE, Kessing LV (2021) Voice analyses using smartphone-based data in patients with bipolar disorder, unaffected relatives and healthy control individuals, and during different affective states. Int J Bipolar Disord 9(1):1\u201313","journal-title":"Int J Bipolar Disord"},{"key":"870_CR6","doi-asserted-by":"publisher","first-page":"351","DOI":"10.21437\/Interspeech.2022-113","volume":"2022","author":"D Feinberg","year":"2022","unstructured":"Feinberg D (2022) Voicelab: Software for fully reproducible automated voice analysis. Proc Interspeech 2022:351\u2013355","journal-title":"Proc Interspeech"},{"key":"870_CR7","doi-asserted-by":"crossref","unstructured":"Feinberg D, Cook O (2021) VoiceLab: Automated reproducible acoustical analysis. https:\/\/github.com\/Voice-Lab\/VoiceLab#voicelab","DOI":"10.31234\/osf.io\/v5uxf"},{"issue":"4","key":"870_CR8","doi-asserted-by":"publisher","first-page":"1738","DOI":"10.1121\/1.399423","volume":"87","author":"H Hermansky","year":"1990","unstructured":"Hermansky H (1990) Perceptual linear predictive (plp) analysis of speech. J Acoust Soc Am 87(4):1738\u20131752","journal-title":"J Acoust Soc Am"},{"issue":"2","key":"870_CR9","doi-asserted-by":"publisher","first-page":"105","DOI":"10.1016\/S0385-8146(12)80192-1","volume":"17","author":"H Hirabayashi","year":"1990","unstructured":"Hirabayashi H, Koshii K, Uno K, Ohgaki H, Nakasone Y, Fujisawa T, Shono N, Hinohara T, Hirabayashi K (1990) Laryngeal epithelial changes on effects of smoking and drinking. Auris Nasus Larynx 17(2):105\u2013114","journal-title":"Auris Nasus Larynx"},{"key":"870_CR10","doi-asserted-by":"crossref","unstructured":"Ioffe S (2006) Probabilistic linear discriminant analysis. In: European Conference on Computer Vision. Springer, pp. 531\u2013542","DOI":"10.1007\/11744085_41"},{"issue":"2","key":"870_CR11","first-page":"87","volume":"57","author":"K Ishihara","year":"1986","unstructured":"Ishihara K, Miyashita A, Inukami M, Fukuda K, Yamazaki K, Miyata H (1986) Results of a Japanese Morningness-Eveningness questionnaire survey. Psychol Res 57(2):87\u201391","journal-title":"Psychol Res"},{"key":"870_CR12","doi-asserted-by":"crossref","unstructured":"Larcher A, Lee KA, Meignier S (2016) An extensible speaker identification sidekit in python. In: 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5095\u20135099. IEEE","DOI":"10.1109\/ICASSP.2016.7472648"},{"key":"870_CR13","unstructured":"Mayuko K, Ryuichi N, Toshio I, Hidenori K et al (2013) Voice tells your body information. Research Report Special Interest Group on MUSic and computer (MUS) 2013(47):1\u20136"},{"key":"870_CR14","unstructured":"Ministry of Health, Labour and Welfare: Overview of 2020 vital statistics monthly report (approximate). https:\/\/www.mhlw.go.jp\/toukei\/saikin\/hw\/jinkou\/geppo\/nengai20\/. Accessed on 12 Nov 2021"},{"key":"870_CR15","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2019.101027","volume":"60","author":"A Nagrani","year":"2020","unstructured":"Nagrani A, Chung JS, Xie W, Zisserman A (2020) VoxCeleb: Large-scale speaker verification in the wild. Comput Speech Lang 60:101027","journal-title":"Comput Speech Lang"},{"issue":"3","key":"870_CR16","doi-asserted-by":"publisher","first-page":"115","DOI":"10.1159\/000266302","volume":"46","author":"G Niedzielsk","year":"1994","unstructured":"Niedzielsk G, Pruszewicz A, \u015awidzi\u0144ski P (1994) Acoustic evaluation of voice in individuals with alcohol addiction. Folia Phoniatr Logop 46(3):115\u2013122","journal-title":"Folia Phoniatr Logop"},{"key":"870_CR17","doi-asserted-by":"crossref","unstructured":"Poorjam AH, Hesaraki S, Safavi S, van Hamme H, Bahari MH (2017) Automatic smoker detection from telephone speech signals. In: International Conference on Speech and Computer. Springer, pp 200\u2013210","DOI":"10.1007\/978-3-319-66429-3_19"},{"key":"870_CR18","unstructured":"Ravanelli M, Parcollet T, Plantinga P, Rouhe A, Cornell S, Lugosch L, Subakan C, Dawalatabad N, Heba A, Zhong J, Chou JC, Yeh SL, Fu SW, Liao CF, Rastorgueva E, Grondin F, Aris W, Na H, Gao Y, Mori RD, Bengio Y (2021) SpeechBrain: A general-purpose speech toolkit. ArXiv:2106.04624"},{"key":"870_CR19","doi-asserted-by":"crossref","unstructured":"Snyder D, Garcia-Romero D, Sell G, Povey D, Khudanpur S (2018) X-vectors: Robust DNN embeddings for speaker recognition. In: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5329\u20135333. IEEE","DOI":"10.1109\/ICASSP.2018.8461375"},{"key":"870_CR20","unstructured":"Sojitra RB (2020) Probabilistic linear discriminant analysis. https:\/\/github.com\/RaviSoji\/plda"},{"key":"870_CR21","unstructured":"Speech Resources Consortium: Atr phoneme balance 503 sentence. http:\/\/research.nii.ac.jp\/src\/ATR503.html. Accessed 12 Nov 2021"},{"issue":"3","key":"870_CR22","doi-asserted-by":"publisher","first-page":"625","DOI":"10.1007\/s10772-020-09726-7","volume":"23","author":"SV Viswanath","year":"2020","unstructured":"Viswanath SV, Swarna K, Prasuna K (2020) An efficient state detection of a person by fusion of acoustic and alcoholic features using various classification algorithms. Int J Speech Technol 23(3):625\u2013632","journal-title":"Int J Speech Technol"},{"issue":"1","key":"870_CR23","doi-asserted-by":"publisher","first-page":"19","DOI":"10.4103\/jlv.JLV_15_18","volume":"8","author":"T Zacharia","year":"2018","unstructured":"Zacharia T, Souza P, Mathew M, Souza G, James J, Baliga M (2018) Effect of circadian cycle on voice: a cross-sectional study with young adults of different chronotypes. J Laryngol Voice 8(1):19\u201323. https:\/\/doi.org\/10.4103\/jlv.JLV_15_18","journal-title":"J Laryngol Voice"}],"container-title":["Artificial Life and Robotics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10015-023-00870-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10015-023-00870-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10015-023-00870-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,7,20]],"date-time":"2023-07-20T06:04:04Z","timestamp":1689833044000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10015-023-00870-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,1]]},"references-count":23,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,8]]}},"alternative-id":["870"],"URL":"https:\/\/doi.org\/10.1007\/s10015-023-00870-2","relation":{},"ISSN":["1433-5298","1614-7456"],"issn-type":[{"value":"1433-5298","type":"print"},{"value":"1614-7456","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,1]]},"assertion":[{"value":"26 March 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 March 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 June 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}