{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T04:27:50Z","timestamp":1760243270974,"version":"build-2065373602"},"reference-count":29,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2014,5,28]],"date-time":"2014-05-28T00:00:00Z","timestamp":1401235200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/3.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>One of the main issues within the field of social robotics is to endow robots with the ability to direct attention to people with whom they are interacting. Different approaches follow bio-inspired mechanisms, merging audio and visual cues to localize a person using multiple sensors. However, most of these fusion mechanisms have been used in fixed systems, such as those used in video-conference rooms, and thus, they may incur difficulties when constrained to the sensors with which a robot can be equipped. Besides, within the scope of interactive autonomous robots, there is a lack in terms of evaluating the benefits of audio-visual attention mechanisms, compared to only audio or visual approaches, in real scenarios. Most of the tests conducted have been within controlled environments, at short distances and\/or with off-line performance measurements. With the goal of demonstrating the benefit of fusing sensory information with a Bayes inference for interactive robotics, this paper presents a system for localizing a person by processing visual and audio data. Moreover, the performance of this system is evaluated and compared via considering the technical limitations of unimodal systems. The experiments show the promise of the proposed approach for the proactive detection and tracking of speakers in a human-robot interactive framework.<\/jats:p>","DOI":"10.3390\/s140609522","type":"journal-article","created":{"date-parts":[[2014,5,28]],"date-time":"2014-05-28T11:10:36Z","timestamp":1401275436000},"page":"9522-9545","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":21,"title":["Audio-Visual Perception System for a Humanoid Robotic Head"],"prefix":"10.3390","volume":"14","author":[{"given":"Raquel","family":"Viciana-Abad","sequence":"first","affiliation":[{"name":"University of Ja\u00e9n, Multimedia and Multimodal Processing Group, Polytechnic School of Linares, University of Ja\u00e9n Alfonso X El Sabio, 28, 23700, Linares, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rebeca","family":"Marfil","sequence":"additional","affiliation":[{"name":"Dpto. Tecnolog\u00eda Electr\u00f3nica, University of M\u00e1laga, Campus de Teatinos - 29071 M\u00e1laga, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jose","family":"Perez-Lorenzo","sequence":"additional","affiliation":[{"name":"University of Ja\u00e9n, Multimedia and Multimodal Processing Group, Polytechnic School of Linares, University of Ja\u00e9n Alfonso X El Sabio, 28, 23700, Linares, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Juan","family":"Bandera","sequence":"additional","affiliation":[{"name":"Dpto. Tecnolog\u00eda Electr\u00f3nica, University of M\u00e1laga, Campus de Teatinos - 29071 M\u00e1laga, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adrian","family":"Romero-Garces","sequence":"additional","affiliation":[{"name":"Dpto. Tecnolog\u00eda Electr\u00f3nica, University of M\u00e1laga, Campus de Teatinos - 29071 M\u00e1laga, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pedro","family":"Reche-Lopez","sequence":"additional","affiliation":[{"name":"University of Ja\u00e9n, Multimedia and Multimodal Processing Group, Polytechnic School of Linares, University of Ja\u00e9n Alfonso X El Sabio, 28, 23700, Linares, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2014,5,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"111","DOI":"10.1177\/153331750401900209","article-title":"Therapeutic robocat for nursing home residents with dementia: Preliminary inquiry","volume":"19","author":"Libin","year":"2004","journal-title":"Am. J. Alzheimers Dis. Other Demen."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1016\/S0921-8890(02)00372-X","article-title":"A survey of socially interactive robots","volume":"42","author":"Fong","year":"2003","journal-title":"Robot. Auton. Syst."},{"key":"ref_3","unstructured":"Koene, A., Mor\u00e9n, J., Trifa, V., and Cheng, G. (2007, January 21\u201324). Gaze shift reflex in a humanoid active vision system. Bielefeld, Germany."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1142\/S0219843608001285","article-title":"Biologically based top-down attention modulation for humanoid interactions","volume":"5","author":"Ude","year":"2008","journal-title":"Int. J. Humanoid Robots"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"699","DOI":"10.1109\/TSMCB.2012.2214477","article-title":"A Bayesian Framework for Active Artificial Perception","volume":"43","author":"Ferreira","year":"2013","journal-title":"IEEE Trans. Cybern."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Nakamura, K., Nakadai, K., Asano, F., and Ince, G. (2011, January 25\u201330). Intelligent sound source localization and its application to multimodal human tracking. San Francisco, CA, USA.","DOI":"10.1109\/IROS.2011.6048166"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1430","DOI":"10.1016\/j.patcog.2006.02.017","article-title":"Pyramid segmentation algorithms revisited","volume":"39","author":"Marfil","year":"2006","journal-title":"Pattern Recognit."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"698","DOI":"10.1016\/j.apacoust.2012.02.002","article-title":"Evaluation of general cross-correlation methods for direction of arrival estimation using two microphones in real environments","volume":"73","author":"Reche","year":"2012","journal-title":"Appl. Acoust."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1940","DOI":"10.1002\/cpe.2816","article-title":"A DDS-based middleware for quality-of-service and high-performance networked robotics","volume":"24","author":"Cruz","year":"2012","journal-title":"Concurr. Comput. Pract. Exp."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"2233","DOI":"10.1016\/j.visres.2010.05.013","article-title":"What and where: A Bayesian inference theory of attention","volume":"50","author":"Chikkerur","year":"2010","journal-title":"Vision Res."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"712","DOI":"10.1016\/j.tins.2004.10.007","article-title":"The Bayesian brain: The role of uncertainty in neural coding and computation","volume":"27","author":"Knill","year":"2004","journal-title":"Trends Neurosci."},{"key":"ref_12","unstructured":"Nakadai, K., Hidai, K., Mizoguchi, H., Okuno, H.G., and Kitano, H. (2011, January 4\u201310). Real-Time Auditory and Visual Multiple-Object Tracking for Humanoids. Seattle, WA, USA."},{"key":"ref_13","first-page":"1727","article-title":"Detection and Separation of Speech Event Using Audio and Video Information Fusion and Its Application to Robust Speech Interface","volume":"2004","author":"Asano","year":"2004","journal-title":"EURASIP J. Adv. Sig. Proc."},{"key":"ref_14","first-page":"2404","article-title":"Robust speech interface based on audio and video information fusion for humanoid HRP-2","volume":"3","author":"Hara","year":"2004","journal-title":"Proc. IROS."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"209","DOI":"10.1527\/tjsai.20.209","article-title":"Dynamic Communication of Humanoid Robot with Multiple People Based on Interaction Distance","volume":"20","author":"Tasaki","year":"2005","journal-title":"Trans. Jpn. Soc. Artif. Intell."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Trifa, V., Koene, A., Moren, J., and Cheng, G. (2007, January 26\u201329). Real-time acoustic source localization in noisy environments for human-robot multimodal interaction. Jeju Island, Korea.","DOI":"10.1109\/ROMAN.2007.4415116"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"172","DOI":"10.1177\/1059712311434662","article-title":"A hierarchical Bayesian framework for multimodal active perception","volume":"20","author":"Ferreira","year":"2012","journal-title":"Adapt. Behav."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Nickel, K., and Stiefelhagen, R. (2007). Fast audio-visual multi-person tracking for a humanoid stereo camera head. Humanoids, 434\u2013441.","DOI":"10.1109\/ICHR.2007.4813906"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"97","DOI":"10.1016\/0010-0285(80)90005-5","article-title":"A feature integration theory of attention","volume":"12","author":"Treisman","year":"1980","journal-title":"Cogn. Psychol."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"202","DOI":"10.3758\/BF03200774","article-title":"Guided search 2.0: A revised model of visual search","volume":"1","author":"Wolfe","year":"1994","journal-title":"Psychon. Bull. Rev."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Klein, D., and Frintrop, S. (2011, January 6\u201313). Center-surround divergence of feature statistics for salient object detection. Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126499"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1167\/9.3.5","article-title":"Saliency, attention, and visual search: An information theoretic approach","volume":"9","author":"Bruce","year":"2009","journal-title":"J. Vision"},{"key":"ref_23","first-page":"547","article-title":"Bayesian Surprise Attracts Human Attention","volume":"19","author":"Itti","year":"2009","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"429","DOI":"10.1006\/ccog.1997.0302","article-title":"Limited capacity of any realizable perceptual system is a sufficient reason for attentive behavior","volume":"6","author":"Tsotsos","year":"1997","journal-title":"Conscious. Cogn."},{"key":"ref_25","first-page":"1","article-title":"A novel biologically inspired attention mechanism for a social robot","volume":"4","author":"Palomino","year":"2011","journal-title":"EURASIP J. Adv. Signal Process"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Marfil, R., Bandera, A., Rodr\u00edguez, J.A., and Sandoval, F. (2009). A novel hierarchical framework for object-based visual attention. Lecture Notes Artif. Intell., 27\u201340.","DOI":"10.1007\/978-3-642-00582-4_3"},{"key":"ref_27","unstructured":"Terrillon, J.-C., and Akamatsu, S. (1999, January 19\u201321). Comparative performance of different chrominance spaces for color segmentation and detection of human faces in complex scene images. Canada."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1950","DOI":"10.1016\/j.sigpro.2011.09.032","article-title":"Multi-source TDOA estimation in reverberant audio using angular spectra and clustering","volume":"92","author":"Blandin","year":"2012","journal-title":"Signal Process."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1257","DOI":"10.1121\/1.4740489","article-title":"A Bayesian inference model for speech localization","volume":"132","author":"Escolano","year":"2012","journal-title":"J. Acoust. Soc. Am."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/14\/6\/9522\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T21:11:56Z","timestamp":1760217116000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/14\/6\/9522"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,5,28]]},"references-count":29,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2014,6]]}},"alternative-id":["s140609522"],"URL":"https:\/\/doi.org\/10.3390\/s140609522","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2014,5,28]]}}}