{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T03:00:48Z","timestamp":1784084448189,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":45,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,5]],"date-time":"2020-10-05T00:00:00Z","timestamp":1601856000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,5]]},"DOI":"10.1145\/3379503.3403535","type":"proceedings-article","created":{"date-parts":[[2020,10,1]],"date-time":"2020-10-01T16:31:59Z","timestamp":1601569919000},"page":"1-9","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Augmenting Conversational Agents with Ambient Acoustic Contexts"],"prefix":"10.1145","author":[{"given":"Chunjong","family":"Park","sequence":"first","affiliation":[{"name":"University of Washington, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chulhong","family":"Min","sequence":"additional","affiliation":[{"name":"Nokia Bell Labs, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sourav","family":"Bhattacharya","sequence":"additional","affiliation":[{"name":"Samsung AI Center Cambridge, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fahim","family":"Kawsar","sequence":"additional","affiliation":[{"name":"Nokia Bell Labs, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,10,5]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2019. Dialogflow. https:\/\/dialogflow.com. Accessed: 2019-04-17.  2019. Dialogflow. https:\/\/dialogflow.com. Accessed: 2019-04-17."},{"key":"e_1_3_2_1_2_1","unstructured":"2019. IBM Bluemix. https:\/\/console.bluemix.net. Accessed: 2019-04-17.  2019. IBM Bluemix. https:\/\/console.bluemix.net. Accessed: 2019-04-17."},{"key":"e_1_3_2_1_3_1","unstructured":"Daniel Adiwardana Minh-Thang Luong David\u00a0R. So Jamie Hall Noah Fiedel Romal Thoppilan Zi Yang Apoorv Kulshreshtha Gaurav Nemade Yifeng Lu and Quoc\u00a0V. Le. 2020. Towards a Human-like Open-Domain Chatbot. arxiv:2001.09977\u00a0[cs.CL]  Daniel Adiwardana Minh-Thang Luong David\u00a0R. So Jamie Hall Noah Fiedel Romal Thoppilan Zi Yang Apoorv Kulshreshtha Gaurav Nemade Yifeng Lu and Quoc\u00a0V. Le. 2020. Towards a Human-like Open-Domain Chatbot. arxiv:2001.09977\u00a0[cs.CL]"},{"key":"e_1_3_2_1_4_1","volume-title":"Proceedings of The 33rd International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol.\u00a048)","author":"Amodei Dario","year":"2016","unstructured":"Dario Amodei , Sundaram Ananthanarayanan , Rishita Anubhai , Jingliang Bai , Eric Battenberg , Carl Case , Jared Casper , Bryan Catanzaro , Qiang Cheng , Guoliang Chen , 2016 . Deep speech 2: End-to-end speech recognition in english and mandarin . In Proceedings of The 33rd International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol.\u00a048) , Maria\u00a0Florina Balcan and Kilian\u00a0Q. Weinberger (Eds.). PMLR, New York, New York, USA, 173\u2013182. Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, 2016. Deep speech 2: End-to-end speech recognition in english and mandarin. In Proceedings of The 33rd International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol.\u00a048), Maria\u00a0Florina Balcan and Kilian\u00a0Q. Weinberger (Eds.). PMLR, New York, New York, USA, 173\u2013182."},{"key":"e_1_3_2_1_5_1","unstructured":"Christopher Baber. 2002. Developing interactive speech technology. Interactive speech technology: Human factors issues in the application of speech input\/output to computers(2002) 1\u201318.  Christopher Baber. 2002. Developing interactive speech technology. Interactive speech technology: Human factors issues in the application of speech input\/output to computers(2002) 1\u201318."},{"key":"e_1_3_2_1_6_1","first-page":"10","volume-title":"Proceedings of the 21st International Conference on Human-Computer Interaction with Mobile Devices and Services","author":"Baldauf Matthias","year":"2019","unstructured":"Matthias Baldauf , Stefan Ribler , and Peter Fr\u00f6hlich . 2019 . Alexa, I\u2019m in Need! Investigating the Potential and Barriers of Voice Assistance Services for Social Work . In Proceedings of the 21st International Conference on Human-Computer Interaction with Mobile Devices and Services ( Taipei, Taiwan) (MobileHCI 2019). Association for Computing Machinery, New York, NY, USA, Article 50, 6\u00a0pages. https:\/\/doi.org\/ 10 .1145\/3338286.3344397 10.1145\/3338286.3344397 Matthias Baldauf, Stefan Ribler, and Peter Fr\u00f6hlich. 2019. Alexa, I\u2019m in Need! Investigating the Potential and Barriers of Voice Assistance Services for Social Work. In Proceedings of the 21st International Conference on Human-Computer Interaction with Mobile Devices and Services (Taipei, Taiwan) (MobileHCI 2019). Association for Computing Machinery, New York, NY, USA, Article 50, 6\u00a0pages. https:\/\/doi.org\/10.1145\/3338286.3344397"},{"key":"e_1_3_2_1_7_1","unstructured":"Jan\u00a0K Chorowski Dzmitry Bahdanau Dmitriy Serdyuk Kyunghyun Cho and Yoshua Bengio. 2015. Attention-based models for speech recognition. In Advances in neural information processing systems. 577\u2013585.  Jan\u00a0K Chorowski Dzmitry Bahdanau Dmitriy Serdyuk Kyunghyun Cho and Yoshua Bengio. 2015. Attention-based models for speech recognition. In Advances in neural information processing systems. 577\u2013585."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2851581.2886425"},{"key":"e_1_3_2_1_9_1","first-page":"3","article-title":"Optimization of RNN-Based Speech Activity Detection","volume":"26","author":"Gelly Gregory","year":"2018","unstructured":"Gregory Gelly and Jean-Luc Gauvain . 2018 . Optimization of RNN-Based Speech Activity Detection . IEEE\/ACM Trans. Audio, Speech and Lang. Proc. 26 , 3 (March 2018), 646\u2013656. https:\/\/doi.org\/10.1109\/TASLP.2017.2769220 10.1109\/TASLP.2017.2769220 Gregory Gelly and Jean-Luc Gauvain. 2018. Optimization of RNN-Based Speech Activity Detection. IEEE\/ACM Trans. Audio, Speech and Lang. Proc. 26, 3 (March 2018), 646\u2013656. https:\/\/doi.org\/10.1109\/TASLP.2017.2769220","journal-title":"IEEE\/ACM Trans. Audio, Speech and Lang. Proc."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952261"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3131895"},{"key":"e_1_3_2_1_12_1","volume-title":"International conference on machine learning. 1764\u20131772","author":"Graves Alex","year":"2014","unstructured":"Alex Graves and Navdeep Jaitly . 2014 . Towards end-to-end speech recognition with recurrent neural networks . In International conference on machine learning. 1764\u20131772 . Alex Graves and Navdeep Jaitly. 2014. Towards end-to-end speech recognition with recurrent neural networks. In International conference on machine learning. 1764\u20131772."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Kun Han Dong Yu and Ivan Tashev. 2014. Speech emotion recognition using deep neural network and extreme learning machine. (2014).  Kun Han Dong Yu and Ivan Tashev. 2014. Speech emotion recognition using deep neural network and extreme learning machine. (2014).","DOI":"10.21437\/Interspeech.2014-57"},{"key":"e_1_3_2_1_14_1","unstructured":"Awni Hannun Carl Case Jared Casper Bryan Catanzaro Greg Diamos Erich Elsen Ryan Prenger Sanjeev Satheesh Shubho Sengupta Adam Coates 2014. Deep speech: Scaling up end-to-end speech recognition. arXiv preprint arXiv:1412.5567(2014).  Awni Hannun Carl Case Jared Casper Bryan Catanzaro Greg Diamos Erich Elsen Ryan Prenger Sanjeev Satheesh Shubho Sengupta Adam Coates 2014. Deep speech: Scaling up end-to-end speech recognition. arXiv preprint arXiv:1412.5567(2014)."},{"key":"e_1_3_2_1_15_1","volume-title":"CNN Architectures for Large-Scale Audio Classification. In International Conference on Acoustics, Speech and Signal Processing (ICASSP). https:\/\/arxiv.org\/abs\/1609","author":"Hershey Shawn","year":"2017","unstructured":"Shawn Hershey , Sourish Chaudhuri , Daniel P.\u00a0W. Ellis , Jort\u00a0 F. Gemmeke , Aren Jansen , Channing Moore , Manoj Plakal , Devin Platt , Rif\u00a0 A. Saurous , Bryan Seybold , Malcolm Slaney , Ron Weiss , and Kevin Wilson . 2017 . CNN Architectures for Large-Scale Audio Classification. In International Conference on Acoustics, Speech and Signal Processing (ICASSP). https:\/\/arxiv.org\/abs\/1609 .09430 Shawn Hershey, Sourish Chaudhuri, Daniel P.\u00a0W. Ellis, Jort\u00a0F. Gemmeke, Aren Jansen, Channing Moore, Manoj Plakal, Devin Platt, Rif\u00a0A. Saurous, Bryan Seybold, Malcolm Slaney, Ron Weiss, and Kevin Wilson. 2017. CNN Architectures for Large-Scale Audio Classification. In International Conference on Acoustics, Speech and Signal Processing (ICASSP). https:\/\/arxiv.org\/abs\/1609.09430"},{"key":"e_1_3_2_1_16_1","volume-title":"Article arXiv:1503.02531 (March","author":"Hinton Geoffrey","year":"2015","unstructured":"Geoffrey Hinton , Oriol Vinyals , and Jeff Dean . 2015. Distilling the Knowledge in a Neural Network. arXiv e-prints , Article arXiv:1503.02531 (March 2015 ), arXiv:1503.02531\u00a0pages. arxiv:1503.02531\u00a0[stat.ML] Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the Knowledge in a Neural Network. arXiv e-prints, Article arXiv:1503.02531 (March 2015), arXiv:1503.02531\u00a0pages. arxiv:1503.02531\u00a0[stat.ML]"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/MPRV.2019.2922907"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173574.3174042"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/MPRV.2018.03367740"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461392"},{"key":"e_1_3_2_1_21_1","first-page":"202","article-title":"Virtual game assistant based on artificial intelligence","volume":"9","author":"Kuhn J","year":"2015","unstructured":"Michael\u00a0 J Kuhn . 2015 . Virtual game assistant based on artificial intelligence . US Patent 9 , 202 ,171. Michael\u00a0J Kuhn. 2015. Virtual game assistant based on artificial intelligence. US Patent 9,202,171.","journal-title":"US Patent"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2750858.2804262"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3242587.3242609"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3242587.3242609"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3314404"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555816.1555834"},{"key":"e_1_3_2_1_27_1","volume-title":"Computers and conversation","author":"Luff Paul","unstructured":"Paul Luff , David Frohlich , and Nigel\u00a0 G Gilbert . 2014. Computers and conversation . Elsevier . Paul Luff, David Frohlich, and Nigel\u00a0G Gilbert. 2014. Computers and conversation. Elsevier."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2858036.2858288"},{"key":"e_1_3_2_1_29_1","unstructured":"Lindsay\u00a0C Page and Hunter Gehlbach. [n.d.]. How an Artificially Intelligent Virtual Assistant Helps Students Navigate the Road to College. ([n.\u00a0d.]) 12.  Lindsay\u00a0C Page and Hunter Gehlbach. [n.d.]. How an Artificially Intelligent Virtual Assistant Helps Students Navigate the Road to College. ([n.\u00a0d.]) 12."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.bushor.2016.03.004"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/SYSMART.2016.7894543"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/PerComW.2013.6529487"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISWC.2012.12"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"crossref","unstructured":"Tara Sainath and Carolina Parada. 2015. Convolutional Neural Networks for Small-Footprint Keyword Spotting. In Interspeech.  Tara Sainath and Carolina Parada. 2015. Convolutional Neural Networks for Small-Footprint Keyword Spotting. In Interspeech.","DOI":"10.21437\/Interspeech.2015-352"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/IranianCEE.2017.7985293"},{"key":"e_1_3_2_1_36_1","volume-title":"Neural Responding Machine for Short-Text Conversation. arXiv:1503.02364 [cs] (March","author":"Shang Lifeng","year":"2015","unstructured":"Lifeng Shang , Zhengdong Lu , and Hang Li. 2015. Neural Responding Machine for Short-Text Conversation. arXiv:1503.02364 [cs] (March 2015 ). http:\/\/arxiv.org\/abs\/1503.02364 arXiv: 1503.02364. Lifeng Shang, Zhengdong Lu, and Hang Li. 2015. Neural Responding Machine for Short-Text Conversation. arXiv:1503.02364 [cs] (March 2015). http:\/\/arxiv.org\/abs\/1503.02364 arXiv: 1503.02364."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1631\/FITEE.1700826"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3338286.3344391"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10916-017-0771-y"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300772"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300772"},{"key":"e_1_3_2_1_42_1","volume-title":"DCASE 2018 Workshop.","author":"Yu Changsong","year":"2018","unstructured":"Changsong Yu , Karim\u00a0Said Barsim , Qiuqiang Kong , and Bin Yang . 2018 . Multi-level attention model for weakly supervised audio classification . In DCASE 2018 Workshop. Changsong Yu, Karim\u00a0Said Barsim, Qiuqiang Kong, and Bin Yang. 2018. Multi-level attention model for weakly supervised audio classification. In DCASE 2018 Workshop."},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3030024.3040201"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7953077"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00368"}],"event":{"name":"MobileHCI '20: 22nd International Conference on Human-Computer Interaction with Mobile Devices and Services","location":"Oldenburg Germany","acronym":"MobileHCI '20","sponsor":["SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["22nd International Conference on Human-Computer Interaction with Mobile Devices and Services"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3379503.3403535","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3379503.3403535","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:02:23Z","timestamp":1750197743000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3379503.3403535"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,5]]},"references-count":45,"alternative-id":["10.1145\/3379503.3403535","10.1145\/3379503"],"URL":"https:\/\/doi.org\/10.1145\/3379503.3403535","relation":{},"subject":[],"published":{"date-parts":[[2020,10,5]]},"assertion":[{"value":"2020-10-05","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}