{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T18:22:08Z","timestamp":1779906128567,"version":"3.53.1"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2017,7,20]],"date-time":"2017-07-20T00:00:00Z","timestamp":1500508800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2017,8,31]]},"abstract":"<jats:p>Editing audio narration using conventional software typically involves many painstaking low-level manipulations. Some state of the art systems allow the editor to work in a text transcript of the narration, and perform select, cut, copy and paste operations directly in the transcript; these operations are then automatically applied to the waveform in a straightforward manner. However, an obvious gap in the text-based interface is the ability to type new words not appearing in the transcript, for example inserting a new word for emphasis or replacing a misspoken word. While high-quality voice synthesizers exist today, the challenge is to synthesize the new word in a voice that matches the rest of the narration. This paper presents a system that can synthesize a new word or short phrase such that it blends seamlessly in the context of the existing narration. Our approach is to use a text to speech synthesizer to say the word in a generic voice, and then use voice conversion to convert it into a voice that matches the narration. Offering a range of degrees of control to the editor, our interface supports fully automatic synthesis, selection among a candidate set of alternative pronunciations, fine control over edit placements and pitch profiles, and even guidance by the editors own voice. The paper presents studies showing that the output of our method is preferred over baseline methods and often indistinguishable from the original voice.<\/jats:p>","DOI":"10.1145\/3072959.3073702","type":"journal-article","created":{"date-parts":[[2017,7,21]],"date-time":"2017-07-21T12:24:07Z","timestamp":1500639847000},"page":"1-13","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":39,"title":["VoCo"],"prefix":"10.1145","volume":"36","author":[{"given":"Zeyu","family":"Jin","sequence":"first","affiliation":[{"name":"Princeton University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gautham J.","family":"Mysore","sequence":"additional","affiliation":[{"name":"Adobe Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stephen","family":"Diverdi","sequence":"additional","affiliation":[{"name":"Adobe Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jingwan","family":"Lu","sequence":"additional","affiliation":[{"name":"Adobe Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Adam","family":"Finkelstein","sequence":"additional","affiliation":[{"name":"Princeton University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2017,7,20]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"Acapela Group. 2016. http:\/\/www.acapela-group.com. (2016). Accessed: 2016-04-10.  Acapela Group. 2016. http:\/\/www.acapela-group.com. (2016). Accessed: 2016-04-10."},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2014.6855137"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185563"},{"key":"e_1_2_2_4_1","unstructured":"Paulus Petrus Gerardus Boersma etal 2002. Praat a system for doing phonetics by computer. Glot international 5 (2002).  Paulus Petrus Gerardus Boersma et al. 2002. Praat a system for doing phonetics by computer. Glot international 5 (2002)."},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/258734.258880"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/778712.778737"},{"key":"e_1_2_2_7_1","first-page":"1859","article-title":"Voice conversion using deep neural networks with layer-wise generative training. Audio, Speech, and Language Processing","volume":"22","author":"Chen Ling-Hui","year":"2014","unstructured":"Ling-Hui Chen , Zhen-Hua Ling , Li-Juan Liu , and Li-Rong Dai . 2014 . Voice conversion using deep neural networks with layer-wise generative training. Audio, Speech, and Language Processing , IEEE\/ACM Transactions on 22 , 12 (2014), 1859 -- 1872 . Ling-Hui Chen, Zhen-Hua Ling, Li-Juan Liu, and Li-Rong Dai. 2014. Voice conversion using deep neural networks with layer-wise generative training. Audio, Speech, and Language Processing, IEEE\/ACM Transactions on 22, 12 (2014), 1859--1872.","journal-title":"IEEE\/ACM Transactions on"},{"key":"e_1_2_2_8_1","volume-title":"Progress in speech synthesis","author":"Conkie Alistair D","unstructured":"Alistair D Conkie and Stephen Isard . 1997. Optimal coupling of diphones . In Progress in speech synthesis . Springer , 293--304. Alistair D Conkie and Stephen Isard. 1997. Optimal coupling of diphones. In Progress in speech synthesis. Springer, 293--304."},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2009.4960478"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2007.366962"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/PROC.1973.9030"},{"key":"e_1_2_2_12_1","first-page":"1617","article-title":"High-Individuality Voice Conversion Based on Concatenative Speech Synthesis. International Journal of Electrical, Computer","volume":"1","author":"Fujii Kei","year":"2007","unstructured":"Kei Fujii , Jun Okawa , and Kaori Suigetsu . 2007 . High-Individuality Voice Conversion Based on Concatenative Speech Synthesis. International Journal of Electrical, Computer , Energetic, Electronic and Communication Engineering 1 , 11 (2007), 1617 -- 1622 . Kei Fujii, Jun Okawa, and Kaori Suigetsu. 2007. High-Individuality Voice Conversion Based on Concatenative Speech Synthesis. International Journal of Electrical, Computer, Energetic, Electronic and Communication Engineering 1, 11 (2007), 1617 -- 1622.","journal-title":"Energetic, Electronic and Communication Engineering"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7471747"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/383259.383295"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1996.541110"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2016.7472761"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1998.674423"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2008.4518514"},{"key":"e_1_2_2_19_1","volume-title":"Fifth ISCA Workshop on Speech Synthesis.","author":"Kominek John","year":"2004","unstructured":"John Kominek and Alan W Black . 2004 . The CMU Arctic speech databases . In Fifth ISCA Workshop on Speech Synthesis. John Kominek and Alan W Black. 2004. The CMU Arctic speech databases. In Fifth ISCA Workshop on Speech Synthesis."},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACRIM.1993.407206"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1661412.1618518"},{"key":"e_1_2_2_22_1","article-title":"HelpingHand","volume":"31","author":"Lu Jingwan","year":"2012","unstructured":"Jingwan Lu , Fisher Yu , Adam Finkelstein , and Stephen DiVerdi . 2012 . HelpingHand : Example-based Stroke Stylization. ACM Trans. Graph. 31 , 4, Article 46 (July 2012), 10 pages. Jingwan Lu, Fisher Yu, Adam Finkelstein, and Stephen DiVerdi. 2012. HelpingHand: Example-based Stroke Stylization. ACM Trans. Graph. 31, 4, Article 46 (July 2012), 10 pages.","journal-title":"Example-based Stroke Stylization. ACM Trans. Graph."},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2461956"},{"key":"e_1_2_2_24_1","volume-title":"Proc. Sound and Music Computing (SMC)","author":"Machado Anderson F","year":"2010","unstructured":"Anderson F Machado and Marcelo Queiroz . 2010 . Voice conversion: A critical survey . Proc. Sound and Music Computing (SMC) (2010), 1--8. Anderson F Machado and Marcelo Queiroz. 2010. Voice conversion: A critical survey. Proc. Sound and Music Computing (SMC) (2010), 1--8."},{"key":"e_1_2_2_25_1","volume-title":"Voice recognition algorithms using mel frequency cepstral coefficient (MFCC) and dynamic time warping (DTW) techniques. arXiv preprint arXiv:1003.4083","author":"Muda Lindasalwa","year":"2010","unstructured":"Lindasalwa Muda , Mumtaj Begam , and Irraivan Elamvazuthi . 2010. Voice recognition algorithms using mel frequency cepstral coefficient (MFCC) and dynamic time warping (DTW) techniques. arXiv preprint arXiv:1003.4083 ( 2010 ). Lindasalwa Muda, Mumtaj Begam, and Irraivan Elamvazuthi. 2010. Voice recognition algorithms using mel frequency cepstral coefficient (MFCC) and dynamic time warping (DTW) techniques. arXiv preprint arXiv:1003.4083 (2010)."},{"key":"e_1_2_2_26_1","volume-title":"Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499","author":"van den Oord Aaron","year":"2016","unstructured":"Aaron van den Oord , Sander Dieleman , Heiga Zen , Karen Simonyan , Oriol Vinyals , Alex Graves , Nal Kalchbrenner , Andrew Senior , and Koray Kavukcuoglu . 2016 . Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 (2016). Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 (2016)."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2807442.2807502"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2642918.2647400"},{"key":"e_1_2_2_29_1","volume-title":"Interspeech","author":"Raj Bhiksha","year":"2010","unstructured":"Bhiksha Raj , Tuomas Virtanen , Sourish Chaudhuri , and Rita Singh . 2010. Non-negative matrix factorization based compensation of music for automatic speech recognition . In Interspeech 2010 . 717--720. Bhiksha Raj, Tuomas Virtanen, Sourish Chaudhuri, and Rita Singh. 2010. Non-negative matrix factorization based compensation of music for automatic speech recognition. In Interspeech 2010. 717--720."},{"key":"e_1_2_2_30_1","volume-title":"EUROSPEECH","author":"Roelands Marc","year":"1993","unstructured":"Marc Roelands and Werner Verhelst . 1993 . Waveform similarity based overlap-add (WSOLA) for time-scale modification of speech: structures and evaluation . In EUROSPEECH 1993. 337--340. Marc Roelands and Werner Verhelst. 1993. Waveform similarity based overlap-add (WSOLA) for time-scale modification of speech: structures and evaluation. In EUROSPEECH 1993. 337--340."},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2501988.2501993"},{"key":"e_1_2_2_32_1","volume-title":"Proceedings of Fonetik","author":"Sj\u00f6lander K\u00e5re","year":"2003","unstructured":"K\u00e5re Sj\u00f6lander . 2003 . An HMM-based system for automatic segmentation and alignment of speech . In Proceedings of Fonetik 2003. 93--96. K\u00e5re Sj\u00f6lander. 2003. An HMM-based system for automatic segmentation and alignment of speech. In Proceedings of Fonetik 2003. 93--96."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015706.1015753"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/89.661472"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511816338"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2007.907344"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2007.367303"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2001.941046"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2013.2251852"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/985692.985759"},{"key":"e_1_2_2_41_1","volume-title":"INTERSPEECH","author":"Wu Zhizheng","year":"2013","unstructured":"Zhizheng Wu , Tuomas Virtanen , Tomi Kinnunen , Engsiong Chng , and Haizhou Li . 2013 . Exemplar-based unit selection for voice conversion utilizing temporal information . In INTERSPEECH 2013. 3057--3061. Zhizheng Wu, Tuomas Virtanen, Tomi Kinnunen, Engsiong Chng, and Haizhou Li. 2013. Exemplar-based unit selection for voice conversion utilizing temporal information. In INTERSPEECH 2013. 3057--3061."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3072959.3073702","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3072959.3073702","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:30:23Z","timestamp":1750217423000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3072959.3073702"}},"subtitle":["text-based insertion and replacement in audio narration"],"short-title":[],"issued":{"date-parts":[[2017,7,20]]},"references-count":41,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2017,8,31]]}},"alternative-id":["10.1145\/3072959.3073702"],"URL":"https:\/\/doi.org\/10.1145\/3072959.3073702","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,7,20]]},"assertion":[{"value":"2017-07-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}