{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,5,14]],"date-time":"2025-05-14T02:31:22Z","timestamp":1747189882291,"version":"3.40.5"},"reference-count":46,"publisher":"World Scientific Pub Co Pte Ltd","issue":"06","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Human. Robot."],"published-print":{"date-parts":[[2020,12]]},"abstract":"<jats:p> Humans make extensive use of auditory cues to interact with other humans, especially in challenging real-world acoustic environments. Multiple distinct acoustic events usually mix together in a complex auditory scene. The ability to separate and localize mixed sound in complex auditory scenes remains a demanding skill for binaural robots. In fact, binaural robots are required to disambiguate and interpret the environmental scene with only two sensors. At the same time, robots that interact with humans should be able to gain insights about the speakers in the environment, such as how many speakers are present and where they are located. For this reason, the speech signal is distinctly important among auditory stimuli commonly found in human-centered acoustic environments. In this paper, we propose a Bayesian method of selectively processing acoustic data that exploits the characteristic amplitude envelope dynamics of human speech to infer the location of speakers in the complex auditory scene. The goal was to demonstrate the effectiveness of this speech-specific temporal dynamics approach. Further, we measure how effective this method is in comparison with more traditional methods based on amplitude detection only. <\/jats:p>","DOI":"10.1142\/s0219843620500231","type":"journal-article","created":{"date-parts":[[2020,12,14]],"date-time":"2020-12-14T07:17:07Z","timestamp":1607930227000},"page":"2050023","source":"Crossref","is-referenced-by-count":1,"title":["Speech Envelope Dynamics for Noise-Robust Auditory Scene Analysis in Robotics"],"prefix":"10.1142","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8535-223X","authenticated-orcid":false,"given":"Francesco","family":"Rea","sequence":"first","affiliation":[{"name":"Robotics Brain and Cognitive Sciences, Istituto Italiano di Tecnologia, Via Enrico Melen, 83, 16152 Genova, GE, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Austin","family":"Kothig","sequence":"additional","affiliation":[{"name":"Social and Intelligent Robotics Research Lab (SIRRL), University of Waterloo, 200 University Avenue West, Waterloo, ON N2L 3G1, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lukas","family":"Grasse","sequence":"additional","affiliation":[{"name":"Department of Neuroscience, University of Lethbridge, 4401 University Drive West, Lethbridge, AB T1K 3M4, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Matthew","family":"Tata","sequence":"additional","affiliation":[{"name":"Department of Neuroscience, University of Lethbridge, 4401 University Drive West, Lethbridge, AB T1K 3M4, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2021,1,21]]},"reference":[{"volume-title":"Auditory Scene Analysis: The Perceptual Organization of Sound","year":"1994","author":"Bregman A. S.","key":"S0219843620500231BIB001"},{"key":"S0219843620500231BIB002","doi-asserted-by":"publisher","DOI":"10.1121\/1.1490592"},{"key":"S0219843620500231BIB003","doi-asserted-by":"publisher","DOI":"10.1121\/1.1510141"},{"key":"S0219843620500231BIB004","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2015.03.003"},{"key":"S0219843620500231BIB005","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2017.07.011"},{"key":"S0219843620500231BIB006","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1976.1162830"},{"key":"S0219843620500231BIB007","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAS.2010.5670137"},{"key":"S0219843620500231BIB008","doi-asserted-by":"publisher","DOI":"10.1163\/016918610X493561"},{"key":"S0219843620500231BIB009","doi-asserted-by":"publisher","DOI":"10.1007\/s10514-012-9316-x"},{"key":"S0219843620500231BIB010","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-03892-6_17"},{"key":"S0219843620500231BIB011","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2006.885251"},{"issue":"11","key":"S0219843620500231BIB012","volume":"2015","author":"Rascon C.","year":"2015","journal-title":"EURASIP J. Audio Speech Music Process."},{"volume-title":"Spatial Hearing: The Psychophysics of Human Sound Localization","year":"1997","author":"Blauert E.","key":"S0219843620500231BIB013"},{"key":"S0219843620500231BIB014","doi-asserted-by":"publisher","DOI":"10.1109\/72.761727"},{"key":"S0219843620500231BIB015","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1003985"},{"key":"S0219843620500231BIB016","doi-asserted-by":"publisher","DOI":"10.1121\/1.4923448"},{"key":"S0219843620500231BIB017","doi-asserted-by":"publisher","DOI":"10.1299\/jbse.6.26"},{"key":"S0219843620500231BIB018","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2013.6696769"},{"key":"S0219843620500231BIB019","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178457"},{"key":"S0219843620500231BIB020","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2001.977176"},{"key":"S0219843620500231BIB021","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2012.6385967"},{"key":"S0219843620500231BIB022","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0186104"},{"key":"S0219843620500231BIB023","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2011.6094825"},{"key":"S0219843620500231BIB024","doi-asserted-by":"publisher","DOI":"10.1109\/MFI.2015.7295835"},{"key":"S0219843620500231BIB025","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2016.7487306"},{"key":"S0219843620500231BIB026","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2016.7487652"},{"key":"S0219843620500231BIB027","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.201400998"},{"key":"S0219843620500231BIB028","doi-asserted-by":"publisher","DOI":"10.1016\/j.neuron.2007.06.004"},{"key":"S0219843620500231BIB029","doi-asserted-by":"publisher","DOI":"10.3389\/fpsyg.2011.00130"},{"key":"S0219843620500231BIB030","doi-asserted-by":"publisher","DOI":"10.1038\/nn.3063"},{"key":"S0219843620500231BIB031","doi-asserted-by":"publisher","DOI":"10.1016\/j.neuron.2012.12.037"},{"key":"S0219843620500231BIB032","doi-asserted-by":"publisher","DOI":"10.1523\/JNEUROSCI.3631-09.2010"},{"key":"S0219843620500231BIB033","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1205381109"},{"key":"S0219843620500231BIB034","doi-asserted-by":"publisher","DOI":"10.1016\/j.bandl.2014.05.003"},{"key":"S0219843620500231BIB035","doi-asserted-by":"publisher","DOI":"10.1016\/j.bandl.2018.12.005"},{"key":"S0219843620500231BIB036","doi-asserted-by":"publisher","DOI":"10.1109\/ROSE.2019.8790411"},{"key":"S0219843620500231BIB037","doi-asserted-by":"publisher","DOI":"10.1145\/1774674.1774683"},{"key":"S0219843620500231BIB040","first-page":"1","volume-title":"Annex C of the SVOS Final Report: Part A: The Auditory Filterbank","volume":"1","author":"Holdsworth J.","year":"1988"},{"volume-title":"Experimental Psychology","year":"1954","author":"Woodworth R. S.","key":"S0219843620500231BIB042"},{"volume-title":"Computer Music: Synthesis, Composition and Performance","year":"1997","author":"Dodge C.","key":"S0219843620500231BIB043"},{"key":"S0219843620500231BIB044","doi-asserted-by":"publisher","DOI":"10.1121\/1.1915985"},{"key":"S0219843620500231BIB045","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"S0219843620500231BIB046","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2655045"},{"key":"S0219843620500231BIB047","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2006.281843"},{"key":"S0219843620500231BIB048","doi-asserted-by":"publisher","DOI":"10.1037\/h0054629"},{"key":"S0219843620500231BIB049","doi-asserted-by":"publisher","DOI":"10.1142\/S0219843615500231"}],"container-title":["International Journal of Humanoid Robotics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0219843620500231","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,3,3]],"date-time":"2021-03-03T11:00:25Z","timestamp":1614769225000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/abs\/10.1142\/S0219843620500231"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,12]]},"references-count":46,"journal-issue":{"issue":"06","published-print":{"date-parts":[[2020,12]]}},"alternative-id":["10.1142\/S0219843620500231"],"URL":"https:\/\/doi.org\/10.1142\/s0219843620500231","relation":{},"ISSN":["0219-8436","1793-6942"],"issn-type":[{"type":"print","value":"0219-8436"},{"type":"electronic","value":"1793-6942"}],"subject":[],"published":{"date-parts":[[2020,12]]}}}