{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,24]],"date-time":"2025-08-24T23:04:45Z","timestamp":1756076685154},"reference-count":22,"publisher":"World Scientific Pub Co Pte Lt","issue":"02","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Semantic Computing"],"published-print":{"date-parts":[[2010,6]]},"abstract":"<jats:p> In this paper, we first show the importance of face-voice correlation for audio-visual person recognition. We propose a simple multimodal fusion technique which preserves the correlation between audio-visual features during speech and evaluate the performance of such a system against audio-only, video-only, and audio-visual systems which use audio and visual features neglecting the interdependency of a person's spoken utterance and the associated facial movements. Experiments performed on the VidTIMIT dataset show that the proposed multimodal fusion scheme has a lower error rate than all other comparison conditions and is more robust against replay attacks. The simplicity of the fusion technique allows for low-complexity designs for a simple low-cost real-time DSP implementation. We then discuss some problems associated with the previously proposed design and, as a solution to those problems, propose two novel classifier designs which provide more flexibility and a convenient way to represent multimodal data where each modality has different characteristics. We also show that these novel classifier designs offer superior performance in terms of both accuracy and robustness. <\/jats:p>","DOI":"10.1142\/s1793351x10000985","type":"journal-article","created":{"date-parts":[[2010,11,2]],"date-time":"2010-11-02T10:33:47Z","timestamp":1288694027000},"page":"155-179","source":"Crossref","is-referenced-by-count":4,"title":["ROBUST MULTIMODAL PERSON RECOGNITION USING LOW-COMPLEXITY AUDIO-VISUAL FEATURE FUSION APPROACHES"],"prefix":"10.1142","volume":"04","author":[{"given":"DHAVAL","family":"SHAH","sequence":"first","affiliation":[{"name":"Signal Analysis and Interpretation Laboratory (SAIL), Ming Hsieh Department of Electrical Engineering, Viterbi School of Engineering, University of Southern California, Los Angeles, CA 90089, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"KYU J.","family":"HAN","sequence":"additional","affiliation":[{"name":"Signal Analysis and Interpretation Laboratory (SAIL), Ming Hsieh Department of Electrical Engineering, Viterbi School of Engineering, University of Southern California, Los Angeles, CA 90089, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"SHRIKANTH S.","family":"NARAYANAN","sequence":"additional","affiliation":[{"name":"Signal Analysis and Interpretation Laboratory (SAIL), Ming Hsieh Department of Electrical Engineering, Viterbi School of Engineering, University of Southern California, Los Angeles, CA 90089, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2011,11,21]]},"reference":[{"key":"rf2","doi-asserted-by":"publisher","DOI":"10.1007\/b117227"},{"key":"rf3","doi-asserted-by":"publisher","DOI":"10.1109\/19.930458"},{"key":"rf4","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2003.1251144"},{"key":"rf5","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2003.818349"},{"key":"rf6","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-84882-254-2"},{"key":"rf7","first-page":"955","volume":"12","author":"Brunelli R.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"rf8","doi-asserted-by":"publisher","DOI":"10.1109\/34.667881"},{"key":"rf9","doi-asserted-by":"publisher","DOI":"10.1109\/2.820041"},{"key":"rf11","doi-asserted-by":"publisher","DOI":"10.1109\/MAES.2006.1703234"},{"key":"rf14","doi-asserted-by":"publisher","DOI":"10.1109\/72.788647"},{"key":"rf15","doi-asserted-by":"publisher","DOI":"10.1016\/j.dsp.2004.05.001"},{"key":"rf16","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2006.886017"},{"key":"rf17","first-page":"179","volume":"2007","author":"Bredin H.","journal-title":"EURASIP J. Appl. Signal Process."},{"key":"rf18","volume-title":"Biometric Person Recognition: Face, Speech, and Fusion","author":"Sanderson C.","year":"2008"},{"key":"rf19","first-page":"1","volume":"2009","author":"Karam W.","journal-title":"EURASIP J. Adv. Signal Process."},{"key":"rf20","doi-asserted-by":"publisher","DOI":"10.1121\/1.1907309"},{"key":"rf21","volume-title":"Speech Perception by Ear and Eye: A Paradigm for Psychological Inquiry","author":"Massaro D. W.","year":"1987"},{"key":"rf22","volume-title":"Issues in Visual and Audio-Visual Speech Processing","author":"Potamios G.","year":"2004"},{"key":"rf23","first-page":"1","volume":"39","author":"Dempster A. P.","journal-title":"Journal of the Royal Statistical Society"},{"key":"rf29","doi-asserted-by":"publisher","DOI":"10.1109\/5.628714"},{"key":"rf30","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000013087.49260.fb"},{"key":"rf31","doi-asserted-by":"publisher","DOI":"10.1016\/0167-6393(95)00009-D"}],"container-title":["International Journal of Semantic Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S1793351X10000985","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,8,7]],"date-time":"2019-08-07T16:06:47Z","timestamp":1565194007000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/abs\/10.1142\/S1793351X10000985"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2010,6]]},"references-count":22,"journal-issue":{"issue":"02","published-online":{"date-parts":[[2011,11,21]]},"published-print":{"date-parts":[[2010,6]]}},"alternative-id":["10.1142\/S1793351X10000985"],"URL":"https:\/\/doi.org\/10.1142\/s1793351x10000985","relation":{},"ISSN":["1793-351X","1793-7108"],"issn-type":[{"value":"1793-351X","type":"print"},{"value":"1793-7108","type":"electronic"}],"subject":[],"published":{"date-parts":[[2010,6]]}}}