{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,15]],"date-time":"2025-07-15T03:16:59Z","timestamp":1752549419724,"version":"3.37.3"},"reference-count":32,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2021,8,28]],"date-time":"2021-08-28T00:00:00Z","timestamp":1630108800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,8,28]],"date-time":"2021-08-28T00:00:00Z","timestamp":1630108800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100007064","name":"Universidad de Ja\u00e9n","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100007064","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"published-print":{"date-parts":[[2022,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Improving the ability to interact through voice with a robot is still a challenge especially in real environments where multiple speakers coexist. This work has evaluated a proposal based on improving the intelligibility of the voice information that feeds an existing ASR service in the network and in conditions similar to those that could occur in a care centre for the elderly. The results indicate the feasibility and improvement of a proposal based on the use of an embedded microphone array and the use of a simple beamforming and masking technique. The system has been evaluated with 12 people and results obtained for time responsiveness indicate that the system would allow natural interaction with voice. It is shown to be necessary to incorporate a system to properly employ the masking algorithm, through the intelligent and stable estimation of the interfering signals. In addition, this approach allows to fix as sources of interest other speakers not located in the vicinity of the robot.\u00a0<\/jats:p>","DOI":"10.1007\/s11042-021-11291-3","type":"journal-article","created":{"date-parts":[[2021,8,28]],"date-time":"2021-08-28T16:02:33Z","timestamp":1630166553000},"page":"3327-3350","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["An audio enhancement system to improve intelligibility for social-awareness in HRI"],"prefix":"10.1007","volume":"81","author":[{"given":"Antonio","family":"Mart\u00ednez-Col\u00f3n","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2545-7229","authenticated-orcid":false,"given":"Raquel","family":"Viciana-Abad","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jose Manuel","family":"Perez-Lorenzo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christine","family":"Evers","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Patrick A.","family":"Naylor","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,8,28]]},"reference":[{"doi-asserted-by":"crossref","unstructured":"Becker E, Le Z, Park K, Lin Y, Makedon F (2009) Event-based experiments in an assistive environment using wireless sensor networks and voice recognition. In Proceedings of the 2nd International Conference on PErvasive Technologies Related to Assistive Environments (PETRA \u201909). Association for Computing Machinery, New York, NY, USA, Article 17, 1-8. https:\/\/doi.org\/10.1145\/1579114.1579131","key":"11291_CR1","DOI":"10.1145\/1579114.1579131"},{"doi-asserted-by":"crossref","unstructured":"Biocca F (1997) The cyborg\u2019s dilemma: embodiment in virtual environments. In Proceedings of Second International Conference on Cognitive Technology Humanizing the Information Age, Japan, pp 12-26. https:\/\/doi.org\/10.1109\/CT.1997.617676","key":"11291_CR2","DOI":"10.1109\/CT.1997.617676"},{"doi-asserted-by":"crossref","unstructured":"Chakrabarty S, Habets EAP (2019) Multi-Speaker DOA estimation using deep convolutional networks trained with noise signals. In IEEE J Sel Top Sign Proces\u00a0vol. 13, no. 1, 8-21. https:\/\/doi.org\/10.1109\/JSTSP.2019.2901664","key":"11291_CR3","DOI":"10.1109\/JSTSP.2019.2901664"},{"doi-asserted-by":"crossref","unstructured":"Chang X, Zhang W, Qian Y, Roux JL, Watanabe S (2020) MIMO-Speech: end-to-end multi-channel multi-speaker speech recognition. In Proceedings of IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pp. 237-244. https:\/\/doi.org\/10.1109\/ASRU46091.2019.9003986","key":"11291_CR4","DOI":"10.1109\/ASRU46091.2019.9003986"},{"unstructured":"DiBiase JH, Silverman HF, Brandstein MS (2001) Microphone arrays: signal processing techniques and applications. M. S. Brandstein and D. Ward, Eds. Springer-Verlag","key":"11291_CR5"},{"doi-asserted-by":"crossref","unstructured":"Evers C, Moore AH, Naylor PA, Sheaffer J, Rafaely B (2015) Bearing-only acoustic tracking of moving speakers for robot audition. In Proceedings of 2015 IEEE International Conference on Digital Signal Processing (DSP), Singapore (Singapur)","key":"11291_CR6","DOI":"10.1109\/ICDSP.2015.7252071"},{"doi-asserted-by":"crossref","unstructured":"Evers C, Naylor PA (2018) Acoustic SLAM. IEEE\/ACM Trans Audio, Speech and Lang Proc 26, 9, 1484-1498. https:\/\/doi.org\/10.1109\/TASLP.2018.2828321","key":"11291_CR7","DOI":"10.1109\/TASLP.2018.2828321"},{"doi-asserted-by":"crossref","unstructured":"Garnerin M, Rossato S, Laurent B (2019) Gender representation in French broadcast corpora and its impact on ASR performance. In: 1st International Workshop on AI for Smart TV Content Production, Access and Delivery (AI4TV 19), ACM, New York, pp 3?9. https:\/\/doi.org\/10.1145\/3347449.3357480","key":"11291_CR8","DOI":"10.1145\/3347449.3357480"},{"doi-asserted-by":"crossref","unstructured":"Griffiths L, Jim C (1982) An alternative approach to linearly constrained adaptive beamforming. IEEE Trans Antennas Propag\u00a030, 27-34. https:\/\/doi.org\/10.1109\/TSP.2010.2051803","key":"11291_CR9","DOI":"10.1109\/TAP.1982.1142739"},{"key":"11291_CR10","volume-title":"Estimation of sound source number and directions under a multi-source environment. In Proceedings of 2009 IEEE\/RSJ Int Conf Intell Robots Syst (IROS 2009)","author":"J Hu","year":"2009","unstructured":"Hu J, Yang C, Wang C (2009) Estimation of sound source number and directions under a multi-source environment. In Proceedings of 2009 IEEE\/RSJ Int Conf Intell Robots Syst (IROS 2009). St, Louis, MO, USA"},{"unstructured":"Jankowski C, Mruthyunjaya V, Lin R (2020) Improved robust ASR for social robots in public spaces.\u00a0https:\/\/arxiv.org\/abs\/2001.04619\u00a0","key":"11291_CR11"},{"doi-asserted-by":"crossref","unstructured":"Kennedy J, Lemaignan S, Montassier C, Lavalade P, Irfan B, Papadopoulos F, Senft E, Belpaeme T (2017) Child speech recognition in human-robot interaction: evaluations and recommendations. In: 12th ACM\/IEEE International Conference on Human-Robot Interaction (HRI), IEEE\/ACM, Vienna, pp 82?90. https:\/\/doi.org\/10.1145\/2909824.3020229","key":"11291_CR12","DOI":"10.1145\/2909824.3020229"},{"unstructured":"Kriegel J, Grabner V, Tuttle-Weidinger L, Ehrenmuller I (2019) Socially Assistive Robots (SAR) in in-patient care for the elderly. Stud Health Technol Inform 260: 178-185. https:\/\/doi.org\/10.3233\/978-1-61499-971-3-178","key":"11291_CR13"},{"doi-asserted-by":"crossref","unstructured":"Lazzeri N, Mazzei D, Cominelli L, Cisternino A, De Rossi D (2018) Designing the mind of a social robot. Appl Sci\u00a08, 302. https:\/\/doi.org\/10.3390\/app8020302","key":"11291_CR14","DOI":"10.3390\/app8020302"},{"issue":"1","key":"11291_CR15","doi-asserted-by":"publisher","first-page":"112","DOI":"10.1109\/TCE.2015.7064118","volume":"61","author":"H Lim","year":"2015","unstructured":"Lim H, Yoo I, Cho Y, Yook D (2015) Speaker localization in noisy environments using steered response voice power. IEEE Trans Consum Electron 61(1):112\u2013118","journal-title":"IEEE Trans Consum Electron"},{"doi-asserted-by":"crossref","unstructured":"Matamoros M, Harbusch K, Paulus D (2018) From commands to goal-based dialogs: A roadmap to achieve natural language interaction in RoboCup@Home. In: Holz D., Genter K., Saad M., von Stryk O. (eds) RoboCup 2018: Robot World Cup XXII. RoboCup 2018. Lect Notes Comput Sci vol 11374. Springer, Cham. https:\/\/doi.org\/10.1007\/978-3-030-27544-0_18","key":"11291_CR16","DOI":"10.1007\/978-3-030-27544-0_18"},{"doi-asserted-by":"crossref","unstructured":"Martinez J et al (2018) Towards a robust robotic assistant for Comprehensive Geriatric Assessment procedures: updating the CLARC system. In Proceedings of 27th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN), IEEE Press, Nanjing, pp. 820-25. https:\/\/doi.org\/10.1109\/ROMAN.2018.8525818","key":"11291_CR17","DOI":"10.1109\/ROMAN.2018.8525818"},{"doi-asserted-by":"crossref","unstructured":"Martinez-Colon A, Perez-Lorenzo JM, Rivas F, Viciana-Abad R, Reche-Lopez P (2018) Attentional mechanism based on a microphone array for embedded devices and a single camera. In Proceedings of the 19th International Workshop of Physical Agents (WAF 2018), November 22-23, Madrid, Spain. https:\/\/doi.org\/10.1007\/978-3-319-99885-5_12","key":"11291_CR18","DOI":"10.1007\/978-3-319-99885-5_12"},{"doi-asserted-by":"crossref","unstructured":"Martinez-Colon A, Viciana-Abad R, Perez-Lorenzo JM, Evers C, Naylor PA (2021) Evaluation of a multi-speaker system for socially assistive HRI in real scenarios. Bergasa, Luis M., Ocana, Manuel, Barea, Rafael, Lopez-Guillen, Elena and Revenga, Pedro (eds.) In Advances in Physical Agents II, WAF 2020\u00a0vol. 1285, Springer, pp 151-166. https:\/\/doi.org\/10.1007\/978-3-030-62579-5_11","key":"11291_CR19","DOI":"10.1007\/978-3-030-62579-5_11"},{"unstructured":"Morgan JP (2017) Time-frequency masking performance for improved intelligibility with microphone arrays. Master Thesis in the College of Engineering at the University of Kentucky","key":"11291_CR20"},{"key":"11291_CR21","doi-asserted-by":"publisher","first-page":"105","DOI":"10.1037\/h0055960","volume":"44","author":"GA Miller","year":"1947","unstructured":"Miller GA (1947) The masking of speech. Psychol Bull 44:105\u2013129. https:\/\/doi.org\/10.1037\/h0055960","journal-title":"Psychol Bull"},{"doi-asserted-by":"crossref","unstructured":"Nikunen J, Diment A, Virtanen T (2018) Separation of moving sound sources using multichannel NMF and acoustic trackings. IEEE\/ACM Trans Audio Speech Lang Process\u00a026, 281-295. https:\/\/doi.org\/10.1109\/TASLP.2017.2774925","key":"11291_CR22","DOI":"10.1109\/TASLP.2017.2774925"},{"doi-asserted-by":"crossref","unstructured":"Okuno HG, Nakadai K, Kim H (2009) Robot audition: missing feature theory approach and active audition. Springer Tracts in Advanced Robotics (14th Conference Robotics Research), 70: 227-244. https:\/\/doi.org\/10.1007\/978-3-642-19457-3_14","key":"11291_CR23","DOI":"10.1007\/978-3-642-19457-3_14"},{"doi-asserted-by":"crossref","unstructured":"Pavlidi D, Puigt M, Griffin A, Mouchtaris A (2012) Real-time multiple sound source localization using a circular microphone array based on single-source confidence measures. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 2625-2628. https:\/\/doi.org\/10.1109\/ICASSP.2012.6288455","key":"11291_CR24","DOI":"10.1109\/ICASSP.2012.6288455"},{"key":"11291_CR25","first-page":"1","volume":"1","author":"C Rascon","year":"2015","unstructured":"Rascon C, Fuentes G, Meza I (2015) Lightweight multi-DOA tracking of mobile speech sources. EURASIP J on Audio, Speech, and Music Processing 1:1\u201316","journal-title":"EURASIP J on Audio, Speech, and Music Processing"},{"key":"11291_CR26","doi-asserted-by":"publisher","first-page":"184","DOI":"10.1016\/j.robot.2017.07.011","volume":"96","author":"C Rascon","year":"2017","unstructured":"Rascon C, Meza I (2017) Localization of sound sources in robotics: A review. Robot Auton Syst 96:184\u2013210","journal-title":"Robot Auton Syst"},{"doi-asserted-by":"crossref","unstructured":"Reche PJ et al (2018) Binaural lateral localization of multiple sources in real environments using a kurtosis-driven split-EM algorithm. Eng Appl Artif Intell 69, 137-146. https:\/\/doi.org\/10.1016\/j.engappai.2017.12.013","key":"11291_CR27","DOI":"10.1016\/j.engappai.2017.12.013"},{"doi-asserted-by":"crossref","unstructured":"Takeda R, Komatani K (2016) Discriminative multiple sound source localization based on deep neural networks using independent location model, In: 2016 IEEE Spoken Language Technology Workshop (SLT), pp. 603-609. https:\/\/doi.org\/10.1109\/SLT.2016.7846325","key":"11291_CR28","DOI":"10.1109\/SLT.2016.7846325"},{"doi-asserted-by":"crossref","unstructured":"Wang D, Chen J (2018) Supervised Speech Separation Based on Deep Learning: An Overview. IEEE\/ACM Trans Audio Speech Lang Process\u00a026: 1702-1726. https:\/\/doi.org\/10.3233\/978-1-61499-971-3-178","key":"11291_CR29","DOI":"10.1109\/TASLP.2018.2842159"},{"doi-asserted-by":"crossref","unstructured":"Valin J, Michaud F, Hadjou B, Rouat J (2004) Localization of simultaneous moving sound sources for mobile robot using a frequency-domain steered beamformer approach. In Proceedings of the IEEE International Conference on Robotics and Automation, 2004 ICRA \u201904, New Orleans, USA","key":"11291_CR30","DOI":"10.1109\/ROBOT.2004.1307286"},{"doi-asserted-by":"crossref","unstructured":"Valin J, Yamamoto S, Rouat J, Michaud F, Nakadai K, Okuno HG (2007) Robust recognition of simultaneous speech by a mobile robot. IEEE Trans Robot\u00a023: 742-752. https:\/\/doi.org\/10.1109\/TRO.2007.900612","key":"11291_CR31","DOI":"10.1109\/TRO.2007.900612"},{"doi-asserted-by":"crossref","unstructured":"Zhuo DB, Cao H (2021) Fast sound source localization based on SRP-PHAT using density peaks clustering. Appl Sci 11, 445. https:\/\/doi.org\/10.3390\/app11010445","key":"11291_CR32","DOI":"10.3390\/app11010445"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-021-11291-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-021-11291-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-021-11291-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,2,21]],"date-time":"2022-02-21T19:40:17Z","timestamp":1645472417000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-021-11291-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,8,28]]},"references-count":32,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,1]]}},"alternative-id":["11291"],"URL":"https:\/\/doi.org\/10.1007\/s11042-021-11291-3","relation":{},"ISSN":["1380-7501","1573-7721"],"issn-type":[{"type":"print","value":"1380-7501"},{"type":"electronic","value":"1573-7721"}],"subject":[],"published":{"date-parts":[[2021,8,28]]},"assertion":[{"value":"6 February 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 May 2021","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 July 2021","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 August 2021","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}