{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,15]],"date-time":"2026-04-15T18:09:31Z","timestamp":1776276571362,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":63,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,4,29]],"date-time":"2022-04-29T00:00:00Z","timestamp":1651190400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,4,29]]},"DOI":"10.1145\/3491102.3517687","type":"proceedings-article","created":{"date-parts":[[2022,4,28]],"date-time":"2022-04-28T16:34:58Z","timestamp":1651163698000},"page":"1-16","source":"Crossref","is-referenced-by-count":5,"title":["Aware: Intuitive Device Activation Using Prosody for Natural Voice Interactions"],"prefix":"10.1145","author":[{"given":"Xinlei","family":"Zhang","sequence":"first","affiliation":[{"name":"Graduate School of Interdisciplinary Information Studies \/ Rekimoto Lab, The University of Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zixiong","family":"Su","sequence":"additional","affiliation":[{"name":"Graduate School of Interdisciplinary Information Studies \/ Rekimoto Lab, The University of Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Rekimoto","sequence":"additional","affiliation":[{"name":"Graduate School of Interdisciplinary Information Studies \/ Rekimoto Lab, The University of Tokyo, Japan and Sony CSL Kyoto, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,4,29]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3379337.3415588"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3469595.3469608"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3311956"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"crossref","unstructured":"Rainer Banse and Klaus\u00a0R Scherer. 1996. Acoustic profiles in vocal emotion expression.Journal of personality and social psychology 70 3(1996) 614.  Rainer Banse and Klaus\u00a0R Scherer. 1996. Acoustic profiles in vocal emotion expression.Journal of personality and social psychology 70 3(1996) 614.","DOI":"10.1037\/0022-3514.70.3.614"},{"key":"e_1_3_2_2_5_1","volume-title":"Subarashii: Encounters in Japanese spoken language education. CALICO journal","author":"Bernstein Jared","year":"1999","unstructured":"Jared Bernstein , Amir Najmi , and Farzad Ehsani . 1999 . Subarashii: Encounters in Japanese spoken language education. CALICO journal (1999), 361\u2013384. Jared Bernstein, Amir Najmi, and Farzad Ehsani. 1999. Subarashii: Encounters in Japanese spoken language education. CALICO journal (1999), 361\u2013384."},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/800250.807503"},{"key":"e_1_3_2_2_7_1","volume-title":"PowerCut and Obfuscator: An Exploration of the Design Space for Privacy-Preserving Interventions for Smart Speakers. In Seventeenth Symposium on Usable Privacy and Security ({SOUPS}","author":"Chandrasekaran Varun","year":"2021","unstructured":"Varun Chandrasekaran , Suman Banerjee , Bilge Mutlu , and Kassem Fawaz . 2021 . PowerCut and Obfuscator: An Exploration of the Design Space for Privacy-Preserving Interventions for Smart Speakers. In Seventeenth Symposium on Usable Privacy and Security ({SOUPS} 2021). 535\u2013552. Varun Chandrasekaran, Suman Banerjee, Bilge Mutlu, and Kassem Fawaz. 2021. PowerCut and Obfuscator: An Exploration of the Design Space for Privacy-Preserving Interventions for Smart Speakers. In Seventeenth Symposium on Usable Privacy and Security ({SOUPS} 2021). 535\u2013552."},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Ailbhe\u00a0N\u00ed Chasaide Irena Yanushevskaya and Christer Gobl. 2017. Voice-to-Affect Mapping: Inferences on Language Voice Baseline Settings.. In INTERSPEECH. 1258\u20131262.  Ailbhe\u00a0N\u00ed Chasaide Irena Yanushevskaya and Christer Gobl. 2017. Voice-to-Affect Mapping: Inferences on Language Voice Baseline Settings.. In INTERSPEECH. 1258\u20131262.","DOI":"10.21437\/Interspeech.2017-1181"},{"key":"e_1_3_2_2_9_1","volume-title":"The sound of sarcasm. Speech communication 50, 5","author":"Cheang S","year":"2008","unstructured":"Henry\u00a0 S Cheang and Marc\u00a0 D Pell . 2008. The sound of sarcasm. Speech communication 50, 5 ( 2008 ), 366\u2013381. Henry\u00a0S Cheang and Marc\u00a0D Pell. 2008. The sound of sarcasm. Speech communication 50, 5 (2008), 366\u2013381."},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3313831.3376304"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF02285283"},{"key":"e_1_3_2_2_12_1","unstructured":"Alice Coucke Alaa Saade Adrien Ball Th\u00e9odore Bluche Alexandre Caulier David Leroy Cl\u00e9ment Doumouro Thibault Gisselbrecht Francesco Caltagirone Thibaut Lavril 2018. Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces. arXiv preprint arXiv:1805.10190(2018).  Alice Coucke Alaa Saade Adrien Ball Th\u00e9odore Bluche Alexandre Caulier David Leroy Cl\u00e9ment Doumouro Thibault Gisselbrecht Francesco Caltagirone Thibaut Lavril 2018. Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces. arXiv preprint arXiv:1805.10190(2018)."},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/SPED.2019.8906584"},{"key":"e_1_3_2_2_14_1","volume-title":"Blendie. In Proceedings of the 5th conference on Designing interactive systems: processes, practices, methods, and techniques. 309\u2013309","author":"Dobson Kelly","year":"2004","unstructured":"Kelly Dobson . 2004 . Blendie. In Proceedings of the 5th conference on Designing interactive systems: processes, practices, methods, and techniques. 309\u2013309 . Kelly Dobson. 2004. Blendie. In Proceedings of the 5th conference on Designing interactive systems: processes, practices, methods, and techniques. 309\u2013309."},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.2478\/popets-2020-0072"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3117811.3117823"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0892-1997(02)00123-6"},{"key":"e_1_3_2_2_18_1","volume-title":"Auditory distance estimation in an open space. Soundscape Semiotics-Localization and Categorization","author":"Fluitt Kim","year":"2014","unstructured":"Kim Fluitt , Timothy Mermagen , and Tomasz Letowski . 2014. Auditory distance estimation in an open space. Soundscape Semiotics-Localization and Categorization ( 2014 ). Kim Fluitt, Timothy Mermagen, and Tomasz Letowski. 2014. Auditory distance estimation in an open space. Soundscape Semiotics-Localization and Categorization (2014)."},{"key":"e_1_3_2_2_19_1","volume-title":"Effnet: An efficient structure for convolutional neural networks. In 2018 25th ieee international conference on image processing (icip)","author":"Freeman Ido","year":"2018","unstructured":"Ido Freeman , Lutz Roese-Koerner , and Anton Kummert . 2018 . Effnet: An efficient structure for convolutional neural networks. In 2018 25th ieee international conference on image processing (icip) . IEEE , 6\u201310. Ido Freeman, Lutz Roese-Koerner, and Anton Kummert. 2018. Effnet: An efficient structure for convolutional neural networks. In 2018 25th ieee international conference on image processing (icip). IEEE, 6\u201310."},{"key":"e_1_3_2_2_20_1","volume-title":"Communicating emotion: The role of prosodic features.Psychological bulletin 97, 3","author":"Frick W","year":"1985","unstructured":"Robert\u00a0 W Frick . 1985. Communicating emotion: The role of prosodic features.Psychological bulletin 97, 3 ( 1985 ), 412. Robert\u00a0W Frick. 1985. Communicating emotion: The role of prosodic features.Psychological bulletin 97, 3 (1985), 412."},{"key":"e_1_3_2_2_21_1","unstructured":"J. Garofolo Lori Lamel W. Fisher Jonathan Fiscus D. Pallett N. Dahlgren and V. Zue. 1992. TIMIT Acoustic-phonetic Continuous Speech Corpus. Linguistic Data Consortium (11 1992).  J. Garofolo Lori Lamel W. Fisher Jonathan Fiscus D. Pallett N. Dahlgren and V. Zue. 1992. TIMIT Acoustic-phonetic Continuous Speech Corpus. Linguistic Data Consortium (11 1992)."},{"key":"e_1_3_2_2_22_1","volume-title":"The role of voice quality in communicating emotion, mood and attitude. Speech communication 40, 1-2","author":"Gobl Christer","year":"2003","unstructured":"Christer Gobl and Ailbhe\u00a0N\u0131 Chasaide . 2003. The role of voice quality in communicating emotion, mood and attitude. Speech communication 40, 1-2 ( 2003 ), 189\u2013212. Christer Gobl and Ailbhe\u00a0N\u0131 Chasaide. 2003. The role of voice quality in communicating emotion, mood and attitude. Speech communication 40, 1-2 (2003), 189\u2013212."},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2004-575"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jml.2016.01.001"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cogdev.2020.100971"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/502348.502372"},{"key":"e_1_3_2_2_27_1","unstructured":"Amazon Inc.[n.d.]. Amazon Echo. https:\/\/www.amazon.com\/smart-home-devices\/b?ie=UTF8&node=9818047011.  Amazon Inc.[n.d.]. Amazon Echo. https:\/\/www.amazon.com\/smart-home-devices\/b?ie=UTF8&node=9818047011."},{"key":"e_1_3_2_2_28_1","unstructured":"Apple Inc.[n.d.]. Siri. https:\/\/www.apple.com\/siri\/.  Apple Inc.[n.d.]. Siri. https:\/\/www.apple.com\/siri\/."},{"key":"e_1_3_2_2_29_1","unstructured":"Google Inc.[n.d.]. Google Nest. https:\/\/store.google.com\/category\/connected_home?.  Google Inc.[n.d.]. Google Nest. https:\/\/store.google.com\/category\/connected_home?."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445169"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2004-240"},{"key":"e_1_3_2_2_32_1","volume-title":"Proc. of Speech Prosody. Citeseer, 883\u2013886","author":"Ishi C","year":"2006","unstructured":"C Ishi , Hiroshi Ishiguro , and Norihiro Hagita . 2006 . Using prosodic and voice quality features for paralinguistic information extraction . In Proc. of Speech Prosody. Citeseer, 883\u2013886 . C Ishi, Hiroshi Ishiguro, and Norihiro Hagita. 2006. Using prosodic and voice quality features for paralinguistic information extraction. In Proc. of Speech Prosody. Citeseer, 883\u2013886."},{"key":"e_1_3_2_2_33_1","volume-title":"Automatic extraction of paralinguistic information using prosodic features related to F0, duration and voice quality. Speech communication 50, 6","author":"Ishi Carlos\u00a0Toshinori","year":"2008","unstructured":"Carlos\u00a0Toshinori Ishi , Hiroshi Ishiguro , and Norihiro Hagita . 2008. Automatic extraction of paralinguistic information using prosodic features related to F0, duration and voice quality. Speech communication 50, 6 ( 2008 ), 531\u2013543. Carlos\u00a0Toshinori Ishi, Hiroshi Ishiguro, and Norihiro Hagita. 2008. Automatic extraction of paralinguistic information using prosodic features related to F0, duration and voice quality. Speech communication 50, 6 (2008), 531\u2013543."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2017.01.011"},{"key":"e_1_3_2_2_35_1","unstructured":"Bjorn Karmann. 2019. Project Alias. https:\/\/bjoernkarmann.dk\/project_alias\/.  Bjorn Karmann. 2019. Project Alias. https:\/\/bjoernkarmann.dk\/project_alias\/."},{"key":"e_1_3_2_2_36_1","unstructured":"Akinobu Lee Tatsuya Kawahara and Kiyohiro Shikano. 2001. Julius\u2014an open source real-time large vocabulary recognition engine. (2001).  Akinobu Lee Tatsuya Kawahara and Kiyohiro Shikano. 2001. Julius\u2014an open source real-time large vocabulary recognition engine. (2001)."},{"key":"e_1_3_2_2_37_1","unstructured":"Matrix. [n.d.]. Matrix Voice. https:\/\/www.matrix.one\/products\/voice.  Matrix. [n.d.]. Matrix Voice. https:\/\/www.matrix.one\/products\/voice."},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3313831.3376479"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3359278"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.2478\/popets-2020-0026"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neubiorev.2013.01.027"},{"key":"e_1_3_2_2_42_1","unstructured":"Jack Mostow 2001. Evaluating tutors that listen: An overview of Project LISTEN.(2001).  Jack Mostow 2001. Evaluating tutors that listen: An overview of Project LISTEN.(2001)."},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/IE.2016.42"},{"key":"e_1_3_2_2_44_1","volume-title":"Computer-assisted telephone interviewing: a general introduction. Telephone survey methodology","author":"Nicholls II","year":"1988","unstructured":"II Nicholls . 1988. WL. Computer-assisted telephone interviewing: a general introduction. Telephone survey methodology . New York : John Wiley & Sons Inc ( 1988 ), 377\u201385. II Nicholls. 1988. WL. Computer-assisted telephone interviewing: a general introduction. Telephone survey methodology. New York: John Wiley & Sons Inc (1988), 377\u201385."},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3404835.3462964"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"crossref","unstructured":"Leonardo Pepino Pablo Riera and Luciana Ferrer. 2021. Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings. arXiv preprint arXiv:2104.03502(2021).  Leonardo Pepino Pablo Riera and Luciana Ferrer. 2021. Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings. arXiv preprint arXiv:2104.03502(2021).","DOI":"10.21437\/Interspeech.2021-703"},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2390256.2390282"},{"key":"e_1_3_2_2_48_1","volume-title":"Considering Wake Gestures for Smart Assistant Use. In Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems. 1\u20138.","author":"Pomykalski Patryk","year":"2020","unstructured":"Patryk Pomykalski , Miko\u0142aj\u00a0 P Wo\u017aniak , Pawe\u0142\u00a0 W Wo\u017aniak , Krzysztof Grudzie\u0144 , Shengdong Zhao , and Andrzej Romanowski . 2020 . Considering Wake Gestures for Smart Assistant Use. In Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems. 1\u20138. Patryk Pomykalski, Miko\u0142aj\u00a0P Wo\u017aniak, Pawe\u0142\u00a0W Wo\u017aniak, Krzysztof Grudzie\u0144, Shengdong Zhao, and Andrzej Romanowski. 2020. Considering Wake Gestures for Smart Assistant Use. In Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems. 1\u20138."},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445687"},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"crossref","unstructured":"Simon Rigoulot Karyn Fish and Marc\u00a0D Pell. 2014. Neural correlates of inferring speaker sincerity from white lies: An event-related potential source localization study. Brain research 1565(2014) 48\u201362.  Simon Rigoulot Karyn Fish and Marc\u00a0D Pell. 2014. Neural correlates of inferring speaker sincerity from white lies: An event-related potential source localization study. Brain research 1565(2014) 48\u201362.","DOI":"10.1016\/j.brainres.2014.04.022"},{"key":"e_1_3_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3239092.3265968"},{"key":"e_1_3_2_2_52_1","volume-title":"15th {USENIX} Symposium on Networked Systems Design and Implementation ({NSDI} 18). 547\u2013560.","author":"Roy Nirupam","unstructured":"Nirupam Roy , Sheng Shen , Haitham Hassanieh , and Romit\u00a0Roy Choudhury . 2018. Inaudible voice commands: The long-range attack and defense . In 15th {USENIX} Symposium on Networked Systems Design and Implementation ({NSDI} 18). 547\u2013560. Nirupam Roy, Sheng Shen, Haitham Hassanieh, and Romit\u00a0Roy Choudhury. 2018. Inaudible voice commands: The long-range attack and defense. In 15th {USENIX} Symposium on Networked Systems Design and Implementation ({NSDI} 18). 547\u2013560."},{"key":"e_1_3_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/2493190.2493244"},{"key":"e_1_3_2_2_54_1","unstructured":"Lea Sch\u00f6nherr Maximilian Golla Thorsten Eisenhofer Jan Wiele Dorothea Kolossa and Thorsten Holz. 2020. Unacceptable where is my privacy? exploring accidental triggers of smart speakers. arXiv preprint arXiv:2008.00508(2020).  Lea Sch\u00f6nherr Maximilian Golla Thorsten Eisenhofer Jan Wiele Dorothea Kolossa and Thorsten Holz. 2020. Unacceptable where is my privacy? exploring accidental triggers of smart speakers. arXiv preprint arXiv:2008.00508(2020)."},{"key":"e_1_3_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3196709.3196772"},{"key":"e_1_3_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461785"},{"key":"e_1_3_2_2_57_1","doi-asserted-by":"crossref","unstructured":"Edwin Simonnet Sahar Ghannay Nathalie Camelin Yannick Est\u00e8ve and Renato De\u00a0Mori. 2017. ASR error management for improving spoken language understanding. arXiv preprint arXiv:1705.09515(2017).  Edwin Simonnet Sahar Ghannay Nathalie Camelin Yannick Est\u00e8ve and Renato De\u00a0Mori. 2017. ASR error management for improving spoken language understanding. arXiv preprint arXiv:1705.09515(2017).","DOI":"10.21437\/Interspeech.2017-1178"},{"key":"e_1_3_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.3026823"},{"key":"e_1_3_2_2_59_1","volume-title":"Prosody and language processing. Language processing","author":"Warren Paul","year":"1999","unstructured":"Paul Warren . 1999. Prosody and language processing. Language processing ( 1999 ), 155\u2013188. Paul Warren. 1999. Prosody and language processing. Language processing (1999), 155\u2013188."},{"key":"e_1_3_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.1445789"},{"key":"e_1_3_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3313831.3376427"},{"key":"e_1_3_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/223904.223952"},{"key":"e_1_3_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2015.05.086"}],"event":{"name":"CHI '22: CHI Conference on Human Factors in Computing Systems","location":"New Orleans LA USA","acronym":"CHI '22","sponsor":["SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["CHI Conference on Human Factors in Computing Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3491102.3517687","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3491102.3517687","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:17:00Z","timestamp":1750191420000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3491102.3517687"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,29]]},"references-count":63,"alternative-id":["10.1145\/3491102.3517687","10.1145\/3491102"],"URL":"https:\/\/doi.org\/10.1145\/3491102.3517687","relation":{},"subject":[],"published":{"date-parts":[[2022,4,29]]}}}