{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,23]],"date-time":"2025-11-23T15:04:15Z","timestamp":1763910255930,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":47,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,4,29]],"date-time":"2022-04-29T00:00:00Z","timestamp":1651190400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"MEDTEQ"},{"name":"Natural Sciences and Research Council (Canada)","award":["CRDPJ 543332-19"],"award-info":[{"award-number":["CRDPJ 543332-19"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,4,29]]},"DOI":"10.1145\/3491102.3501871","type":"proceedings-article","created":{"date-parts":[[2022,4,28]],"date-time":"2022-04-28T17:09:36Z","timestamp":1651165776000},"page":"1-11","source":"Crossref","is-referenced-by-count":1,"title":["The Sound of Hallucinations: Toward a more convincing emulation of internalized voices"],"prefix":"10.1145","author":[{"given":"Hyejin","family":"Lee","sequence":"first","affiliation":[{"name":"Department of Electrical and Computer Engineering, McGill University, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ruixi","family":"Jiang","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, McGill University, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongjae","family":"Yoo","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, McGill University, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Max","family":"Henry","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, McGill University, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jeremy R.","family":"Cooperstock","sequence":"additional","affiliation":[{"name":"Centre for Intelligent Machines, McGill University, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,4,29]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Voice-hearing and personification: characterizing social qualities of auditory verbal hallucinations in early psychosis. Schizophrenia bulletin 47, 1","author":"Alderson-Day Ben","year":"2021","unstructured":"Ben Alderson-Day , Angela Woods , Peter Moseley , Stephanie Common , Felicity Deamer , Guy Dodgson , and Charles Fernyhough . 2021. Voice-hearing and personification: characterizing social qualities of auditory verbal hallucinations in early psychosis. Schizophrenia bulletin 47, 1 ( 2021 ), 228\u2013236. Ben Alderson-Day, Angela Woods, Peter Moseley, Stephanie Common, Felicity Deamer, Guy Dodgson, and Charles Fernyhough. 2021. Voice-hearing and personification: characterizing social qualities of auditory verbal hallucinations in early psychosis. Schizophrenia bulletin 47, 1 (2021), 228\u2013236."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"crossref","unstructured":"Kirsteen\u00a0M Aldrich Elizabeth\u00a0J Hellier and Judy Edworthy. 2009. What determines auditory similarity? The effect of stimulus group and methodology. Quarterly journal of experimental psychology 62 1(2009) 63\u201383.  Kirsteen\u00a0M Aldrich Elizabeth\u00a0J Hellier and Judy Edworthy. 2009. What determines auditory similarity? The effect of stimulus group and methodology. Quarterly journal of experimental psychology 62 1(2009) 63\u201383.","DOI":"10.1080\/17470210701814451"},{"key":"e_1_3_2_2_3_1","unstructured":"Sercan Arik Gregory Diamos Andrew Gibiansky John Miller Kainan Peng Wei Ping Jonathan Raiman and Yanqi Zhou. 2017. Deep Voice 2: Multi-Speaker Neural Text-to-Speech. arxiv:1705.08947\u00a0[cs.CL]  Sercan Arik Gregory Diamos Andrew Gibiansky John Miller Kainan Peng Wei Ping Jonathan Raiman and Yanqi Zhou. 2017. Deep Voice 2: Multi-Speaker Neural Text-to-Speech. arxiv:1705.08947\u00a0[cs.CL]"},{"key":"e_1_3_2_2_4_1","unstructured":"AV Voice Changer 2021. AV Voice Chaner Software Diamond. https:\/\/www.audio4fun.com\/voice-changer.htm.  AV Voice Changer 2021. AV Voice Chaner Software Diamond. https:\/\/www.audio4fun.com\/voice-changer.htm."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.17743\/jaes.2018.0023"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"crossref","unstructured":"Alice Baird Emilia Parada-Cabaleiro Simone Hantke Felix Burkhardt Nicholas Cummins and Bj\u00f6rn Schuller. 2018. The perception and analysis of the likeability and human likeness of synthesized speech. (2018).  Alice Baird Emilia Parada-Cabaleiro Simone Hantke Felix Burkhardt Nicholas Cummins and Bj\u00f6rn Schuller. 2018. The perception and analysis of the likeability and human likeness of synthesized speech. (2018).","DOI":"10.21437\/Interspeech.2018-1093"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICACCI.2013.6637331"},{"key":"e_1_3_2_2_8_1","volume-title":"Proc. 43rd Annual Conference of the German Acoustical Society","author":"Berger Stephanie","year":"2017","unstructured":"Stephanie Berger , Oliver Niebuhr , and Benno Peters . 2017 . Winning over an audience\u2013A perception-based analysis of prosodic features of charismatic speech . In Proc. 43rd Annual Conference of the German Acoustical Society , Kiel, Germany. 1454\u20131457. Stephanie Berger, Oliver Niebuhr, and Benno Peters. 2017. Winning over an audience\u2013A perception-based analysis of prosodic features of charismatic speech. In Proc. 43rd Annual Conference of the German Acoustical Society, Kiel, Germany. 1454\u20131457."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jvoice.2003.12.004"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1017\/S003329171600115X"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/S2215-0366(17)30427-3"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2006-275"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"crossref","unstructured":"Mireia Diez Luk\u00e1s Burget and Pavel Matejka. 2018. Speaker Diarization based on Bayesian HMM with Eigenvoice Priors.. In Odyssey. 147\u2013154.  Mireia Diez Luk\u00e1s Burget and Pavel Matejka. 2018. Speaker Diarization based on Bayesian HMM with Eigenvoice Priors.. In Odyssey. 147\u2013154.","DOI":"10.21437\/Odyssey.2018-21"},{"key":"e_1_3_2_2_14_1","volume-title":"Virtual reality therapy for refractory auditory verbal hallucinations in schizophrenia: a pilot clinical trial. Schizophrenia research 197","author":"du Sert Olivier\u00a0Percie","year":"2018","unstructured":"Olivier\u00a0Percie du Sert , St\u00e9phane Potvin , Olivier Lipp , Laura Dellazizzo , M\u00e9lanie Laurelli , Richard Breton , Pierre Lalonde , Kingsada Phraxayavong , Kieron O\u2019Connor , Jean-Fran\u00e7ois Pelletier , 2018. Virtual reality therapy for refractory auditory verbal hallucinations in schizophrenia: a pilot clinical trial. Schizophrenia research 197 ( 2018 ), 176\u2013181. Olivier\u00a0Percie du Sert, St\u00e9phane Potvin, Olivier Lipp, Laura Dellazizzo, M\u00e9lanie Laurelli, Richard Breton, Pierre Lalonde, Kingsada Phraxayavong, Kieron O\u2019Connor, Jean-Fran\u00e7ois Pelletier, 2018. Virtual reality therapy for refractory auditory verbal hallucinations in schizophrenia: a pilot clinical trial. Schizophrenia research 197 (2018), 176\u2013181."},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1177\/002383098102400408"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.427148"},{"key":"e_1_3_2_2_17_1","volume-title":"Emotional processing of fear: exposure to corrective information.Psychological bulletin 99, 1","author":"Foa B","year":"1986","unstructured":"Edna\u00a0 B Foa and Michael\u00a0 J Kozak . 1986. Emotional processing of fear: exposure to corrective information.Psychological bulletin 99, 1 ( 1986 ), 20. Edna\u00a0B Foa and Michael\u00a0J Kozak. 1986. Emotional processing of fear: exposure to corrective information.Psychological bulletin 99, 1 (1986), 20."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0892-1997(88)80024-9"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1159\/000261923"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jvoice.2006.07.004"},{"key":"e_1_3_2_2_21_1","volume-title":"Speaker Embeddings Incorporating Acoustic Conditions for Diarization. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 7129\u20137133","author":"Higuchi Yosuke","year":"2020","unstructured":"Yosuke Higuchi , Masayuki Suzuki , and Gakuto Kurata . 2020 . Speaker Embeddings Incorporating Acoustic Conditions for Diarization. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 7129\u20137133 . https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9054273 Yosuke Higuchi, Masayuki Suzuki, and Gakuto Kurata. 2020. Speaker Embeddings Incorporating Acoustic Conditions for Diarization. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 7129\u20137133. https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9054273"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"crossref","unstructured":"Guang Hua Jiwu Huang Yun\u00a0Q Shi Jonathan Goh and Vrizlynn\u00a0LL Thing. 2016. Twenty years of digital audio watermarking\u2014a comprehensive review. Signal processing 128(2016) 222\u2013242.  Guang Hua Jiwu Huang Yun\u00a0Q Shi Jonathan Goh and Vrizlynn\u00a0LL Thing. 2016. Twenty years of digital audio watermarking\u2014a comprehensive review. Signal processing 128(2016) 222\u2013242.","DOI":"10.1016\/j.sigpro.2016.04.005"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2013-107"},{"key":"e_1_3_2_2_24_1","unstructured":"Ye Jia Yu Zhang Ron\u00a0J Weiss Quan Wang Jonathan Shen Fei Ren Zhifeng Chen Patrick Nguyen Ruoming Pang Ignacio\u00a0Lopez Moreno 2018. Transfer learning from speaker verification to multispeaker text-to-speech synthesis. arXiv preprint arXiv:1806.04558(2018).  Ye Jia Yu Zhang Ron\u00a0J Weiss Quan Wang Jonathan Shen Fei Ren Zhifeng Chen Patrick Nguyen Ruoming Pang Ignacio\u00a0Lopez Moreno 2018. Transfer learning from speaker verification to multispeaker text-to-speech synthesis. arXiv preprint arXiv:1806.04558(2018)."},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1998.674423"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/89.876308"},{"key":"e_1_3_2_2_27_1","volume-title":"The Human Takes It All: Humanlike Synthesized Voices Are Perceived as Less Eerie and More Likable. Evidence From a Subjective Ratings Study. Frontiers in neurorobotics 14","author":"K\u00fchne Katharina","year":"2020","unstructured":"Katharina K\u00fchne , Martin\u00a0 H Fischer , and Yuefang Zhou . 2020. The Human Takes It All: Humanlike Synthesized Voices Are Perceived as Less Eerie and More Likable. Evidence From a Subjective Ratings Study. Frontiers in neurorobotics 14 ( 2020 ), 105. Katharina K\u00fchne, Martin\u00a0H Fischer, and Yuefang Zhou. 2020. The Human Takes It All: Humanlike Synthesized Voices Are Perceived as Less Eerie and More Likable. Evidence From a Subjective Ratings Study. Frontiers in neurorobotics 14 (2020), 105."},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAU.1973.1162507"},{"key":"e_1_3_2_2_29_1","volume-title":"Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426(2018).","author":"McInnes Leland","year":"2018","unstructured":"Leland McInnes , John Healy , and James Melville . 2018 . Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426(2018). Leland McInnes, John Healy, and James Melville. 2018. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426(2018)."},{"key":"e_1_3_2_2_30_1","unstructured":"MorphVOX Pro 2021. MorphVOX Pro voice changer. https:\/\/screamingbee.com\/.  MorphVOX Pro 2021. MorphVOX Pro voice changer. https:\/\/screamingbee.com\/."},{"key":"e_1_3_2_2_31_1","volume-title":"Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones. Speech communication 9, 5-6","author":"Moulines Eric","year":"1990","unstructured":"Eric Moulines and Francis Charpentier . 1990. Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones. Speech communication 9, 5-6 ( 1990 ), 453\u2013467. Eric Moulines and Francis Charpentier. 1990. Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones. Speech communication 9, 5-6 (1990), 453\u2013467."},{"key":"e_1_3_2_2_32_1","volume-title":"Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499(2016).","author":"van\u00a0den Oord Aaron","year":"2016","unstructured":"Aaron van\u00a0den Oord , Sander Dieleman , Heiga Zen , Karen Simonyan , Oriol Vinyals , Alex Graves , Nal Kalchbrenner , Andrew Senior , and Koray Kavukcuoglu . 2016 . Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499(2016). Aaron van\u00a0den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499(2016)."},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/HAPTIC.2006.1627122"},{"volume-title":"Voice in social interaction. Vol.\u00a05","author":"Pittam Jeff","key":"e_1_3_2_2_34_1","unstructured":"Jeff Pittam . 1994. Voice in social interaction. Vol.\u00a05 . Sage . Jeff Pittam. 1994. Voice in social interaction. Vol.\u00a05. Sage."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASPAA.1995.482995"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.3390\/jcm9092748"},{"key":"e_1_3_2_2_37_1","first-page":"10","article-title":"Modeling techniques for virtual acoustics","volume":"45","author":"Savioja Lauri","year":"1999","unstructured":"Lauri Savioja . 1999 . Modeling techniques for virtual acoustics . Simulation 45 , 10 (1999), 10 . Lauri Savioja. 1999. Modeling techniques for virtual acoustics. Simulation 45, 10 (1999), 10.","journal-title":"Simulation"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1016\/0092-6566(73)90030-5"},{"key":"e_1_3_2_2_39_1","volume-title":"An overview of voice conversion and its challenges: From statistical modeling to deep learning","author":"Sisman Berrak","year":"2020","unstructured":"Berrak Sisman , Junichi Yamagishi , Simon King , and Haizhou Li. 2020. An overview of voice conversion and its challenges: From statistical modeling to deep learning . IEEE\/ACM Transactions on Audio, Speech, and Language Processing ( 2020 ). Berrak Sisman, Junichi Yamagishi, Simon King, and Haizhou Li. 2020. An overview of voice conversion and its challenges: From statistical modeling to deep learning. IEEE\/ACM Transactions on Audio, Speech, and Language Processing (2020)."},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.physbeh.2005.10.013"},{"key":"e_1_3_2_2_41_1","unstructured":"Johan Sundberg and RT Sataloff. 2005. Vocal tract resonance. (2005).  Johan Sundberg and RT Sataloff. 2005. Vocal tract resonance. (2005)."},{"key":"e_1_3_2_2_42_1","unstructured":"Christophe Veaux Junichi Yamagishi Kirsten MacDonald 2017. Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit. (2017).  Christophe Veaux Junichi Yamagishi Kirsten MacDonald 2017. Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit. (2017)."},{"key":"e_1_3_2_2_43_1","unstructured":"Voicemod 2021. Voicemod: Free Real-time Voice Changer. https:\/\/www.voicemod.net\/.  Voicemod 2021. Voicemod: Free Real-time Voice Changer. https:\/\/www.voicemod.net\/."},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8462665"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"crossref","unstructured":"Thomas Ward Rachel Lister Miriam Fornells-Ambrojo Mar Rus-Calafell Clementine\u00a0J Edwards Conan O\u2019Brien Tom\u00a0KJ Craig and Philippa Garety. 2021. The role of characterisation in everyday voice engagement and AVATAR therapy dialogue. Psychological medicine(2021) 1\u20138.  Thomas Ward Rachel Lister Miriam Fornells-Ambrojo Mar Rus-Calafell Clementine\u00a0J Edwards Conan O\u2019Brien Tom\u00a0KJ Craig and Philippa Garety. 2021. The role of characterisation in everyday voice engagement and AVATAR therapy dialogue. Psychological medicine(2021) 1\u20138.","DOI":"10.1017\/S0033291721000659"},{"key":"e_1_3_2_2_46_1","volume-title":"2012 3rd International Conference on System Science, Engineering Design and Manufacturing Informatization, Vol.\u00a01. IEEE, 255\u2013260","author":"Lu","year":"2012","unstructured":"Lu Xiao-chun, Yin Jun-xun, and Hu Wei-ping. 2012 . A text-independent speaker recognition system based on probabilistic principle component analysis . In 2012 3rd International Conference on System Science, Engineering Design and Manufacturing Informatization, Vol.\u00a01. IEEE, 255\u2013260 . Lu Xiao-chun, Yin Jun-xun, and Hu Wei-ping. 2012. A text-independent speaker recognition system based on probabilistic principle component analysis. In 2012 3rd International Conference on System Science, Engineering Design and Manufacturing Informatization, Vol.\u00a01. IEEE, 255\u2013260."},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"crossref","unstructured":"Yu Zhang Ron\u00a0J Weiss Heiga Zen Yonghui Wu Zhifeng Chen RJ Skerry-Ryan Ye Jia Andrew Rosenberg and Bhuvana Ramabhadran. 2019. Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning. arXiv preprint arXiv:1907.04448(2019).  Yu Zhang Ron\u00a0J Weiss Heiga Zen Yonghui Wu Zhifeng Chen RJ Skerry-Ryan Ye Jia Andrew Rosenberg and Bhuvana Ramabhadran. 2019. Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning. arXiv preprint arXiv:1907.04448(2019).","DOI":"10.21437\/Interspeech.2019-2668"}],"event":{"name":"CHI '22: CHI Conference on Human Factors in Computing Systems","sponsor":["SIGCHI ACM Special Interest Group on Computer-Human Interaction"],"location":"New Orleans LA USA","acronym":"CHI '22"},"container-title":["CHI Conference on Human Factors in Computing Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3491102.3501871","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3491102.3501871","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:53Z","timestamp":1750188653000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3491102.3501871"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,29]]},"references-count":47,"alternative-id":["10.1145\/3491102.3501871","10.1145\/3491102"],"URL":"https:\/\/doi.org\/10.1145\/3491102.3501871","relation":{},"subject":[],"published":{"date-parts":[[2022,4,29]]}}}