{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,26]],"date-time":"2026-03-26T17:24:10Z","timestamp":1774545850718,"version":"3.50.1"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2017,9,11]],"date-time":"2017-09-11T00:00:00Z","timestamp":1505088000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Ministry of Science, ICT","award":["DGIST Research and Development Program (CPS Global center)"],"award-info":[{"award-number":["DGIST Research and Development Program (CPS Global center)"]}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS-1319302 and IIS-1521722"],"award-info":[{"award-number":["CNS-1319302 and IIS-1521722"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2017,9,11]]},"abstract":"<jats:p>Distant emotion recognition (DER) extends the application of speech emotion recognition to the very challenging situation that is determined by variable speaker to microphone distances. The performance of conventional emotion recognition systems degrades dramatically as soon as the microphone is moved away from the mouth of the speaker. This is due to a broad variety of effects such as background noise, feature distortion with distance, overlapping speech from other speakers, and reverberation. This paper presents a novel solution for DER, addressing the key challenges by identification and deletion of features from consideration which are significantly distorted by distance, creating a novel, called Emo2vec, feature modeling and overlapping speech filtering technique, and the use of an LSTM classifier to capture the temporal dynamics of speech states found in emotions. A comprehensive evaluation is conducted on two acted datasets (with artificially generated distance effect) as well as on a new emotional dataset of spontaneous family discussions with audio recorded from multiple microphones placed in different distances. Our solution achieves an average 91.6%, 90.1% and 89.5% accuracy for emotion happy, angry and sad, respectively, across various distances which is more than a 16% increase on average in accuracy compared to the best baseline method.<\/jats:p>","DOI":"10.1145\/3130961","type":"journal-article","created":{"date-parts":[[2017,9,11]],"date-time":"2017-09-11T12:12:26Z","timestamp":1505131946000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Distant Emotion Recognition"],"prefix":"10.1145","volume":"1","author":[{"given":"Asif","family":"Salekin","sequence":"first","affiliation":[{"name":"University of Virginia, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zeya","family":"Chen","sequence":"additional","affiliation":[{"name":"University of Virginia, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mohsin Y.","family":"Ahmed","sequence":"additional","affiliation":[{"name":"University of Virginia, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"John","family":"Lach","sequence":"additional","affiliation":[{"name":"University of Virginia, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Donna","family":"Metz","sequence":"additional","affiliation":[{"name":"University of Southern California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kayla","family":"De La Haye","sequence":"additional","affiliation":[{"name":"University of Southern California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Brooke","family":"Bell","sequence":"additional","affiliation":[{"name":"University of Southern California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"John A.","family":"Stankovic","sequence":"additional","affiliation":[{"name":"University of Virginia, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,9,11]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2017. Google cloud speech API. https:\/\/cloud.google.com\/speech\/. (10 Feb 2017).  2017. Google cloud speech API. https:\/\/cloud.google.com\/speech\/. (10 Feb 2017)."},{"key":"e_1_2_1_2_1","doi-asserted-by":"crossref","unstructured":"2017. Spaces in New Homes. goo.gl\/1z3oVs. (10 Feb 2017).  2017. Spaces in New Homes. goo.gl\/1z3oVs. (10 Feb 2017).","DOI":"10.1016\/S0262-1762(19)30131-2"},{"key":"e_1_2_1_3_1","unstructured":"2017. spectral noise gating algorithm. http:\/\/tinyurl.com\/yard8oe. (1 Jan 2017).  2017. spectral noise gating algorithm. http:\/\/tinyurl.com\/yard8oe. (1 Jan 2017)."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-012-9368-5"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2013.07.002"},{"key":"e_1_2_1_6_1","volume-title":"American Society for Engineering Education (ASEE) Zone Conference Proceedings. 1--7.","author":"Bachu RG","year":"2008"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00779-009-0263-2"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCSLP.2014.6936627"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.19173\/irrodl.v13i4.1234"},{"key":"e_1_2_1_10_1","unstructured":"Christine Evers. 2010. Blind dereverberation of speech from moving and stationary speakers using sequential Monte Carlo methods. (2010).  Christine Evers. 2010. Blind dereverberation of speech from moving and stationary speakers using sequential Monte Carlo methods. (2010)."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1049\/iet-spr:20070046"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874246"},{"key":"e_1_2_1_13_1","doi-asserted-by":"crossref","first-page":"249","DOI":"10.21437\/Interspeech.2011-53","article-title":"Analysis of i-vector Length Normalization in Speaker Recognition Systems","volume":"2011","author":"Garcia-Romero Daniel","year":"2011","journal-title":"Interspeech"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10803-009-0862-9"},{"key":"e_1_2_1_15_1","unstructured":"Yoav Goldberg and Omer Levy. 2014. word2vec explained: Deriving mikolov et al.\u2019s negative-sampling word-embedding method. arXiv preprint arXiv:1402.3722 (2014).  Yoav Goldberg and Omer Levy. 2014. word2vec explained: Deriving mikolov et al.\u2019s negative-sampling word-embedding method. arXiv preprint arXiv:1402.3722 (2014)."},{"key":"e_1_2_1_16_1","volume-title":"Acoustics, speech and signal processing (icassp)","author":"Graves Alex","year":"2013"},{"key":"e_1_2_1_17_1","unstructured":"Awni Hannun Carl Case Jared Casper Bryan Catanzaro Greg Diamos Erich Elsen Ryan Prenger Sanjeev Satheesh Shubho Sengupta Adam Coates etal 2014. Deep speech: Scaling up end-to-end speech recognition. arXiv preprint arXiv:1412.5567 (2014).  Awni Hannun Carl Case Jared Casper Bryan Catanzaro Greg Diamos Erich Elsen Ryan Prenger Sanjeev Satheesh Shubho Sengupta Adam Coates et al. 2014. Deep speech: Scaling up end-to-end speech recognition. arXiv preprint arXiv:1412.5567 (2014)."},{"key":"e_1_2_1_18_1","volume-title":"Machine Audition: Principles, Algorithms and Systems. IGI Global, Hershey PA","author":"Haq S.","year":"2010"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/SSP.2007.4301262"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-85099-1_18"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU.2011.6163922"},{"key":"e_1_2_1_23_1","volume-title":"Proceedings. (ICASSP\u201901)","volume":"1","author":"Krishnamachari Kasturi Rangan","year":"2001"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU.2013.6707732"},{"key":"e_1_2_1_25_1","unstructured":"Sungbok Lee Serdar Yildirim Abe Kazemzadeh and Shrikanth Narayanan. 2005. An articulatory study of emotional speech production.. In Interspeech. 497--500.  Sungbok Lee Serdar Yildirim Abe Kazemzadeh and Shrikanth Narayanan. 2005. An articulatory study of emotional speech production.. In Interspeech. 497--500."},{"key":"e_1_2_1_26_1","unstructured":"ZHOU Lian. 2015. Exploration of the Working Principle and Application of Word2vec. Sci-Tech Information Development 8 Economy 2 (2015) 145--148.  ZHOU Lian. 2015. Exploration of the Working Principle and Application of Word2vec. Sci-Tech Information Development 8 Economy 2 (2015) 145--148."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/VNIS.1989.98737"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/RADIOELEK.2016.7477362"},{"key":"e_1_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Mary L McHugh. 2012. Interrater reliability: the kappa statistic. Biochemia medica 22 3 (2012) 276--282.  Mary L McHugh. 2012. Interrater reliability: the kappa statistic. Biochemia medica 22 3 (2012) 276--282.","DOI":"10.11613\/BM.2012.031"},{"key":"e_1_2_1_30_1","first-page":"3","article-title":"Recurrent neural network based language model","volume":"2","author":"Mikolov Tomas","year":"2010","journal-title":"Interspeech"},{"key":"e_1_2_1_31_1","unstructured":"Tomas Mikolov Ilya Sutskever Kai Chen Greg S Corrado and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems. 3111--3119.   Tomas Mikolov Ilya Sutskever Kai Chen Greg S Corrado and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems. 3111--3119."},{"key":"e_1_2_1_32_1","doi-asserted-by":"crossref","unstructured":"Stephanie Pancoast and Murat Akbacak. 2012. Bag-of-Audio-Words Approach for Multimedia Event Classification.. In Interspeech. 2105--2108.  Stephanie Pancoast and Murat Akbacak. 2012. Bag-of-Audio-Words Approach for Multimedia Event Classification.. In Interspeech. 2105--2108.","DOI":"10.21437\/Interspeech.2012-561"},{"key":"e_1_2_1_33_1","unstructured":"Razvan Pascanu Tomas Mikolov and Yoshua Bengio. 2013. On the difficulty of training recurrent neural networks. ICML (3) 28 (2013) 1310--1318.   Razvan Pascanu Tomas Mikolov and Yoshua Bengio. 2013. On the difficulty of training recurrent neural networks. ICML (3) 28 (2013) 1310--1318."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1147\/sj.393.0705"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1016\/S1071-5819(02)00141-6"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2011.05.002"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11235-011-9624-z"},{"key":"e_1_2_1_38_1","unstructured":"Rajib Rana. 2016. Emotion Classification from Noisy Speech-A Deep Learning Approach. arXiv preprint arXiv:1603.05901 (2016).  Rajib Rana. 2016. Emotion Classification from Noisy Speech-A Deep Learning Approach. arXiv preprint arXiv:1603.05901 (2016)."},{"key":"e_1_2_1_39_1","doi-asserted-by":"crossref","unstructured":"Shourabh Rawat Peter F Schulam Susanne Burger Duo Ding Yipei Wang and Florian Metze. 2013. Robust audio-codebooks for large-scale event detection in consumer videos. (2013).  Shourabh Rawat Peter F Schulam Susanne Burger Duo Ding Yipei Wang and Florian Metze. 2013. Robust audio-codebooks for large-scale event detection in consumer videos. (2013).","DOI":"10.21437\/Interspeech.2013-654"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/HSCMA.2014.6843274"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCSLP.2014.6936711"},{"key":"e_1_2_1_42_1","unstructured":"Xin Rong. 2014. word2vec parameter learning explained. arXiv preprint arXiv:1411.2738 (2014).  Xin Rong. 2014. word2vec parameter learning explained. arXiv preprint arXiv:1411.2738 (2014)."},{"key":"e_1_2_1_43_1","doi-asserted-by":"crossref","unstructured":"Melissa Ryan Janice Murray and Ted Ruffman. 2009. Aging and the perception of emotion: Processing vocal expressions alone and with faces. Experimental aging research 36 1 (2009) 1--22.  Melissa Ryan Janice Murray and Ted Ruffman. 2009. Aging and the perception of emotion: Processing vocal expressions alone and with faces. Experimental aging research 36 1 (2009) 1--22.","DOI":"10.1080\/03610730903418372"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2655045"},{"key":"e_1_2_1_45_1","volume-title":"DAVE: Detecting Agitated Vocal Events. In The Second IEEE\/ACM International Conference on Connected Health: Applications, Systems and Engineering Technologies (CHASE). IEEE\/ACM.","author":"Salekin Asif","year":"2017"},{"key":"e_1_2_1_46_1","volume-title":"Workshop on emotion and computing.","author":"Schroder M","year":"2006"},{"key":"e_1_2_1_47_1","volume-title":"The INTERSPEECH 2013 computational paralinguistics challenge: social signals, conflict, emotion, autism.","author":"Schuller Bj\u00f6rn","year":"2013"},{"key":"e_1_2_1_48_1","volume-title":"Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on. IEEE, 2834--2838","author":"Shokouhi Navid","year":"2013"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2670313"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2010.1132"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7472669"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/T-AFFC.2010.16"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/BigData.Congress.2014.59"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2014.6854660"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2006.262865"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2008.52"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/ChinaSIP.2015.7230458"},{"key":"e_1_2_1_58_1","article-title":"The Research of Speech Emotion Recognition Based on Gaussian Mixture Model. In Applied Mechanics and Materials, Vol. 668","author":"Zhang Wan Li","year":"2014","journal-title":"Trans Tech Publ, 1126--1129."},{"key":"e_1_2_1_59_1","volume-title":"Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on. IEEE, 5755--5759","author":"Zhang Yu","year":"2016"},{"key":"e_1_2_1_60_1","volume-title":"The Observer XT: A tool for the integration and synchronization of multimodal signals. Behavior research methods 41, 3","author":"Zimmerman Patrick H","year":"2009"}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3130961","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3130961","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3130961","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T02:13:20Z","timestamp":1750212800000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3130961"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,9,11]]},"references-count":60,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2017,9,11]]}},"alternative-id":["10.1145\/3130961"],"URL":"https:\/\/doi.org\/10.1145\/3130961","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,9,11]]},"assertion":[{"value":"2017-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-09-11","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}