{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T18:39:24Z","timestamp":1776883164861,"version":"3.51.2"},"reference-count":43,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2022,1,1]],"date-time":"2022-01-01T00:00:00Z","timestamp":1640995200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,1,13]],"date-time":"2022-01-13T00:00:00Z","timestamp":1642032000000},"content-version":"vor","delay-in-days":12,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"published-print":{"date-parts":[[2022,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper analyses the performance of different types of Deep Neural Networks to jointly estimate age and identify gender from speech, to be applied in Interactive Voice Response systems available in call centres. Deep Neural Networks are used, because they have recently demonstrated discriminative and representation capabilities in a wide range of applications, including speech processing problems based on feature extraction and selection. Networks with different sizes are analysed to obtain information on how performance depends on the network architecture and the number of free parameters. The speech corpus used for the experiments is Mozilla\u2019s Common Voice dataset, an open and crowdsourced speech corpus. The results are really good for gender classification, independently of the type of neural network, but improve with the network size. Regarding the classification by age groups, the combination of convolutional neural networks and temporal neural networks seems to be the best option among the analysed, and again, the larger the size of the network, the better the results. The results are promising for use in IVR systems, with the best systems achieving a gender identification error of less than 2% and a classification error by age group of less than 20%.<\/jats:p>","DOI":"10.1007\/s11042-021-11614-4","type":"journal-article","created":{"date-parts":[[2022,1,13]],"date-time":"2022-01-13T19:03:50Z","timestamp":1642100630000},"page":"3535-3552","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":50,"title":["Age group classification and gender recognition from speech with temporal convolutional neural networks"],"prefix":"10.1007","volume":"81","author":[{"given":"H\u00e9ctor A.","family":"S\u00e1nchez-Hevia","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Roberto","family":"Gil-Pita","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Manuel","family":"Utrilla-Manso","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3073-3278","authenticated-orcid":false,"given":"Manuel","family":"Rosa-Zurera","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,1,13]]},"reference":[{"key":"11614_CR1","unstructured":"Abadi M, Agarwal A, Barham P, et al (2015) TensorFlow: large-scale machine learning on heterogeneous systems. http:\/\/tensorflow.org\/. Software available from tensorflow.org"},{"issue":"10","key":"11614_CR2","doi-asserted-by":"publisher","first-page":"1533","DOI":"10.1109\/TASLP.2014.2339736","volume":"22","author":"O Abdel-Hamid","year":"2014","unstructured":"Abdel-Hamid O, Abdel-Rahman M, Jiang H, Deng L, Penn G, Yu D (2014) Convolutional neural network for speech recognition. IEEE\/ACM Transactions on Audio, Speech, and Language Processing 22(10):1533\u20131545","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"11614_CR3","doi-asserted-by":"crossref","unstructured":"Badshah A, Ahmad J, Rahim N, Baik S (2017) Speech emotion recognition from spectrograms with deep convolutional neural network. In: 2017 International conference on platform technology and service (PlatCon), pp 1\u20135","DOI":"10.1109\/PlatCon.2017.7883728"},{"key":"11614_CR4","doi-asserted-by":"crossref","unstructured":"Bahari M, McLaren M, Van Leeuwen D, et al (2012) Age estimation from telephone speech using i-vectors. In: Proceedings of Interspeech 2012. Portland, USA","DOI":"10.21437\/Interspeech.2012-169"},{"key":"11614_CR5","doi-asserted-by":"crossref","unstructured":"Bhat C, Mithum B, Saxena V, Kulkarni V, Kopparapu S (2013) Deploying usable speech enabled ivr systems for mass use. In: 2013 IEEE international conference on human computer interaction (ICHCI), pp 1\u20135","DOI":"10.1109\/ICHCI-IEEE.2013.6887794"},{"key":"11614_CR6","doi-asserted-by":"crossref","unstructured":"Cakir E, Adavanne S, Parascandolo G, Drossos K, Virtanen T (2017) Convolutional recurrent neural networks for bird audio detection. In: 2017 25th European signal processing conference (EUSIPCO), pp 1744\u20131748","DOI":"10.23919\/EUSIPCO.2017.8081508"},{"key":"11614_CR7","doi-asserted-by":"crossref","unstructured":"Cho K, van Merrienboer B, Gulcehre C, Bahdanau D, Bougares F, Schwenk H, Bengio Y (2014) Learning phrase representations using rnn encoder-decoder for statistical machine translation. arXiv:1406.1078","DOI":"10.3115\/v1\/D14-1179"},{"key":"11614_CR8","unstructured":"Chollet F, et al (2015) Keras. https:\/\/keras.io"},{"issue":"3","key":"11614_CR9","first-page":"551","volume":"20","author":"M Couper","year":"2004","unstructured":"Couper M, Singer E, Tourangeau R (2004) Does voice matter? An interactive voice response (IVR) experiment. Journal of Official Statistics 20(3):551\u2013570","journal-title":"Journal of Official Statistics"},{"key":"11614_CR10","doi-asserted-by":"crossref","unstructured":"Devillers L, Vidrascu L (2006) Real-life emotions detection with lexical and paralinguistic cues on human-human call center dialogs. In: INTERSPEECH 2006. International Speech Communication Association, pp 801\u2013804.","DOI":"10.21437\/Interspeech.2006-275"},{"key":"11614_CR11","doi-asserted-by":"crossref","unstructured":"Gao Y, Liu Y, Zhang H, Li Z, Zhu Y, Lin H, Yang M (2020) Estimating GPU memory consumption of deep learning models. In: Proceedings of the 28th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering. ACM, pp 1342\u20131352","DOI":"10.1145\/3368089.3417050"},{"key":"11614_CR12","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1016\/S0167-6393(97)00040-X","volume":"23","author":"A Gorin","year":"1997","unstructured":"Gorin A, Riccardi G, Wright J (1997) How may I help you? Speech Communication 23:113\u2013127","journal-title":"Speech Communication"},{"issue":"8","key":"11614_CR13","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter S, Schmidhuber J (1997) Long short term memory. Neural Computation 9(8):1735\u20131780","journal-title":"Neural Computation"},{"key":"11614_CR14","doi-asserted-by":"publisher","first-page":"20231","DOI":"10.1007\/s11042-017-4646-5","volume":"76","author":"J Huang","year":"2017","unstructured":"Huang J, Li B, Zhu J, Chen J (2017) Age classification with deep learning face representation. Multimedia Tools and Applications 76:20231\u201320247","journal-title":"Multimedia Tools and Applications"},{"key":"11614_CR15","doi-asserted-by":"publisher","first-page":"21603","DOI":"10.1007\/s11042-020-08843-4","volume":"79","author":"M Ilyas","year":"2020","unstructured":"Ilyas M, Othmani A, Nait-ali A (2020) Auditory perception based system for age classification and estimation using dynamic frequency sound. Multimedia Tools and Applications 79:21603\u201331626","journal-title":"Multimedia Tools and Applications"},{"key":"11614_CR16","doi-asserted-by":"publisher","first-page":"372","DOI":"10.1016\/j.ress.2019.01.006","volume":"185","author":"C Jinglong","year":"2019","unstructured":"Jinglong C, Hongjie J, Yanhong C, Qian L (2019) Gated recurrent unit based recurrent neural network for remaining useful life prediction of nonlinear deterioration process. Reliability Engineering and System Safety 185:372\u2013382","journal-title":"Reliability Engineering and System Safety"},{"key":"11614_CR17","doi-asserted-by":"crossref","unstructured":"Kalluri SB, Vijayasenan D, Ganapathy S (2019) A deep neural network based end to end model for joint height and age estimation from short duration speech. In: 2019 IEEE International conference on acoustics, speech and signal processing (ICASSP 2007). IEEE, pp 6580\u20136584","DOI":"10.1109\/ICASSP.2019.8683397"},{"key":"11614_CR18","doi-asserted-by":"crossref","unstructured":"Lea C, Vidal R, Reiter A, Hager GD (2016) Temporal convolutional networks: a unified approach to action segmentation. In: European conference on computer vision. Springer, pp 47\u201354","DOI":"10.1007\/978-3-319-49409-8_7"},{"key":"11614_CR19","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"323","author":"Y LeCun","year":"2015","unstructured":"LeCun Y, Bengio Y, Hinton G (2015) Deep learning. Nature 323:436\u2013444","journal-title":"Nature"},{"key":"11614_CR20","doi-asserted-by":"crossref","unstructured":"Mehrbod N, Grilo A, Zutshi A (2018) Caller-agent pairing in call centers using machine learning techniques with imbalanced data. In: 2018 IEEE International conference on engineering, technology and innovation (ICE\/ITMC). IEEE, pp 1\u20136","DOI":"10.1109\/ICE.2018.8436314"},{"key":"11614_CR21","doi-asserted-by":"crossref","unstructured":"Metze F, Ajmera J, Englert R, Bub U, et al (2007) Comparison of four approaches to age and gender recognition for telephone applications. In: 2007 IEEE International conference on acoustics, speech and signal processing (ICASSP 2007), vol 4, pp IV\u20131089","DOI":"10.1109\/ICASSP.2007.367263"},{"key":"11614_CR22","doi-asserted-by":"crossref","unstructured":"Minematsu N, Sekiguchi M, Hirose K (2002) Automatic estimation of one\u2019s age with his\/her speech based upon acoustic modeling techniques of speakers. In: 2002 IEEE International conference on acoustics, speech, and signal processing (ICASSP 2002), vol 1, pp I\u2013137","DOI":"10.1109\/ICASSP.2002.1005695"},{"key":"11614_CR23","unstructured":"Mohino-Herranz I, Garc\u00eda-G\u00f3mez J, Utrilla-Manso M, Rosa-Zurera M (2018) Precision maximization in anger detection in interactive voice response systems. In: 145th convention of the audio engineering society, paper number, pp 10090"},{"key":"11614_CR24","doi-asserted-by":"crossref","unstructured":"Mubarak E, Shahid T, Mustafa M (2020) Does gender and accent of voice matter?: an interactive voice response (ivr) experiment. In: Proceedings of the 2020 international conference on information and communication technologies and development. ACM Digital Library, pp 739\u2013746","DOI":"10.1145\/3392561.3397588"},{"key":"11614_CR25","doi-asserted-by":"crossref","unstructured":"Neumann M, Vu NT (2017) Attentive convolutional neural network based speech emotion recognition: a study on the impact of input features, signal length, and acted speech. arXiv preprint arXiv:1706.00612","DOI":"10.21437\/Interspeech.2017-917"},{"key":"11614_CR26","doi-asserted-by":"crossref","unstructured":"Pandey A, Wang D (2019) Tcnn: temporal convolutional neural network for real-time speech enhancement in the time domain. In: 2019 IEEE international conference on acoustics, speech and signal processing (ICASSP 2019), pp 6875\u20136879","DOI":"10.1109\/ICASSP.2019.8683634"},{"key":"11614_CR27","doi-asserted-by":"crossref","unstructured":"Pappas D, Androutsopoulos I, Papageorgiou H (2015) Anger detection in call center dialogues. In: 2015 6th IEEE international conference on cognitive infocommunications (CogInfoCom), pp 139\u2013144","DOI":"10.1109\/CogInfoCom.2015.7390579"},{"key":"11614_CR28","doi-asserted-by":"crossref","unstructured":"Park SR, Lee JW (2017) A fully convolutional neural network for speech enhancement. In: Proc. Interspeech, pp 1993\u20131997","DOI":"10.21437\/Interspeech.2017-1465"},{"issue":"3","key":"11614_CR29","doi-asserted-by":"publisher","first-page":"127","DOI":"10.1007\/BF02478291","volume":"9","author":"W Pitts","year":"1947","unstructured":"Pitts W, McCulloch W (1947) How we know universals the perception of auditory and visual forms. Bull Math Biophys 9(3):127\u2013147","journal-title":"Bull Math Biophys"},{"key":"11614_CR30","doi-asserted-by":"crossref","unstructured":"Ranjan S, Hansen JH (2017) Improved gender independent speaker recognition using convolutional neural network based bottleneck features. In: Proceedings of Interspeech, pp 1009\u20131013","DOI":"10.21437\/Interspeech.2017-1182"},{"key":"11614_CR31","first-page":"533","volume":"521","author":"Learning representations by back-propagating errors","year":"1986","unstructured":"Learning representations by back-propagating errors (1986) Rumelhart, D., al. Nature 521:533\u2013536","journal-title":"Nature"},{"key":"11614_CR32","doi-asserted-by":"crossref","unstructured":"S\u00e1nchez-Hevia H, Gil-Pita R, Utrilla-Manso M, Rosa-Zurera M (2019) Convolutional-recurrent neural network for age an gender prediction from speech. In: 2019 signal processing symposium, krakow (Poland). IEEE, pp 246\u2013249","DOI":"10.1109\/SPS.2019.8881961"},{"key":"11614_CR33","doi-asserted-by":"crossref","unstructured":"S\u00e1nchez-Hevia H, Gil-Pita R, Utrilla-Manso M, Rosa-Zurera M (2020) Age and gender recognition from speech using deep neural networks. In: Advances in Physical Agents II. Proceedings of the 21st International Workshop of Physical Agents (WAF 2020). Advances in Intelligent Systems and Computing Series. Springer Nature Switzerland, pp 332\u2013344","DOI":"10.1007\/978-3-030-62579-5_23"},{"issue":"105596","key":"11614_CR34","first-page":"1","volume":"194","author":"S Sengupta","year":"2020","unstructured":"Sengupta S, Basak S, Saikia P, Sayak P, Tsalavoutis V, Atiah F, Ravi V, Peters A (2020) A review of deep learning with special emphasis on architectures, applications and recent trends. Knowledge-Based Systems 194(105596):1\u201333","journal-title":"Knowledge-Based Systems"},{"key":"11614_CR35","doi-asserted-by":"crossref","unstructured":"Ghahremani P, Nidadavolu PN, Chen N, Villalba J, Povey D, Khudanpur S, Dehak N (2018) End-to-end deep neural network age estimation. In: Proceedings of the 19th annual conference of the international speech communication association, INTERSPEECH 2018. ISCA, pp 277\u2013281","DOI":"10.21437\/Interspeech.2018-2015"},{"key":"11614_CR36","doi-asserted-by":"crossref","unstructured":"Markitantov M, Verkholyak O (2019) Automatic recognition of speaker age and gender based on deep neural networks. In: Speech and computer, LNAI, vol 11658. Springer Nature, pp 327\u2013336","DOI":"10.1007\/978-3-030-26061-3_34"},{"key":"11614_CR37","doi-asserted-by":"crossref","unstructured":"Singh R, Raj B, Baker J (2016) Short-term analysis for estimating physical parameters of speakers. In: 2016 4th international conference on biometrics and forensics (IWBF). IEEE, pp 1\u20136","DOI":"10.1109\/IWBF.2016.7449696"},{"key":"11614_CR38","doi-asserted-by":"crossref","unstructured":"Tsang K, Wong K, Kang Y (2020) Age estimation in short speech utterances based on lstm recurrent neural networks. Toronto Working Papers in Linguistics(TWPL) 42:1\u201310","DOI":"10.33137\/twpl.v42i1.33149"},{"key":"11614_CR39","doi-asserted-by":"crossref","unstructured":"Vidrascu L, Devillers L (2006) Real-life emotion representation and detection in call centers data. In: International conference on affective computing and intelligent interaction. Springer, pp 739\u2013746","DOI":"10.1007\/11573548_95"},{"key":"11614_CR40","unstructured":"Wang M, Wang X (2010) Study on the workforce scheduling and routing strategies of heterogeneous agents in call centers. In: Advances in economics, business and management research, vol 159 (Fifth International Conference on Economic and Business Management). Atlantic Press, pp 577\u2013583"},{"key":"11614_CR41","doi-asserted-by":"crossref","unstructured":"Xu Y, Kong Q, Wang W, Plumbley M (2018) Large-scale weakly supervised audio classification using gated convolutional neural network. In: 2018 IEEE International conference on acoustics, speech and signal processing (ICASSP 2018), pp 121\u2013125","DOI":"10.1109\/ICASSP.2018.8461975"},{"key":"11614_CR42","doi-asserted-by":"publisher","first-page":"22524","DOI":"10.1109\/ACCESS.2018.2816163","volume":"6","author":"R Zazo","year":"2018","unstructured":"Zazo R, Nidadavolu P, Chen N, Gonzalez-Rodriguez J, Dehak N (2018) Age estimation in short speech utterances based on lstm recurrent neural networks. IEEE Access 6:22524\u201322530","journal-title":"IEEE Access"},{"issue":"11","key":"11614_CR43","doi-asserted-by":"publisher","first-page":"3212","DOI":"10.1109\/TNNLS.2018.2876865","volume":"30","author":"ZQ Zhao","year":"2019","unstructured":"Zhao ZQ, Zheng P, Xu ST, Wu X (2019) Object detection with deep learning: a review. IEEE Transactios on Neural Networks and Learning Systems 30(11):3212\u20133232","journal-title":"IEEE Transactios on Neural Networks and Learning Systems"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-021-11614-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-021-11614-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-021-11614-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,2,21]],"date-time":"2022-02-21T19:20:20Z","timestamp":1645471220000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-021-11614-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1]]},"references-count":43,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,1]]}},"alternative-id":["11614"],"URL":"https:\/\/doi.org\/10.1007\/s11042-021-11614-4","relation":{},"ISSN":["1380-7501","1573-7721"],"issn-type":[{"value":"1380-7501","type":"print"},{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1]]},"assertion":[{"value":"31 January 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 May 2021","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 September 2021","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 January 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}