{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,5]],"date-time":"2026-02-05T06:02:40Z","timestamp":1770271360106,"version":"3.49.0"},"publisher-location":"New York, NY, USA","edition-number":"1","reference-count":152,"publisher":"ACM","isbn-type":[{"value":"9781450387200","type":"print"}],"license":[{"start":{"date-parts":[[2021,9,10]],"date-time":"2021-09-10T00:00:00Z","timestamp":1631232000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,9,10]]},"DOI":"10.1145\/3477322.3477329","type":"book-chapter","created":{"date-parts":[[2021,10,3]],"date-time":"2021-10-03T05:54:32Z","timestamp":1633240472000},"page":"173-212","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["Building and Designing Expressive Speech Synthesis"],"prefix":"10.1145","author":[{"given":"Matthew P.","family":"Aylett","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Leigh","family":"Clark","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Benjamin R.","family":"Cowan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ilaria","family":"Torre","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,2]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3234695.3236344"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/1776334.1776384"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-1113"},{"key":"e_1_3_2_1_4_1","volume-title":"Speech Prosody 2010-Fifth International Conference.","author":"Andersson S.","unstructured":"S. Andersson , K. Georgila , D. Traum , M. Aylett , and R. A. Clark . 2010. Prediction and realisation of conversational characteristics by utilising spontaneous speech for unit selection . In Speech Prosody 2010-Fifth International Conference. S. Andersson, K. Georgila, D. Traum, M. Aylett, and R. A. Clark. 2010. Prediction and realisation of conversational characteristics by utilising spontaneous speech for unit selection. In Speech Prosody 2010-Fifth International Conference."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2696454.2696464"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.5555\/2900423.2900546"},{"key":"e_1_3_2_1_7_1","article-title":"Realistic transformation of facial and vocal smiles in real-time audiovisual streams","author":"Arias P.","year":"2018","unstructured":"P. Arias , C. Soladie , O. Bouafif , A. Robel , R. Seguier , and J.-J. Aucouturier . 2018 . Realistic transformation of facial and vocal smiles in real-time audiovisual streams . IEEE Transactions on Affective Computing. P. Arias, C. Soladie, O. Bouafif, A. Robel, R. Seguier, and J.-J. Aucouturier. 2018. Realistic transformation of facial and vocal smiles in real-time audiovisual streams. IEEE Transactions on Affective Computing.","journal-title":"IEEE Transactions on Affective Computing."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/3327546.3327667"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-1945"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3236112.3236171"},{"key":"e_1_3_2_1_11_1","volume-title":"Eighth ISCA Workshop on Speech Synthesis.","author":"Aylett M. P.","unstructured":"M. P. Aylett , B. Potard , and C. J. Pidcock . 2013. Expressive speech synthesis: Synthesising ambiguity . In Eighth ISCA Workshop on Speech Synthesis. M. P. Aylett, B. Potard, and C. J. Pidcock. 2013. Expressive speech synthesis: Synthesising ambiguity. In Eighth ISCA Workshop on Speech Synthesis."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2559206.2578868"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2017.2763134"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290607.3310422"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3342775.3342806"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12369-010-0082-7"},{"key":"e_1_3_2_1_17_1","unstructured":"R. Barthes. 1977. Image\u2013Music\u2013Text. Macmillan. R. Barthes. 1977. Image\u2013Music\u2013Text . Macmillan."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/2390470.2390488"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1093\/iwc\/iwv029"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1080\/03740463.2005.10416087"},{"key":"e_1_3_2_1_21_1","volume-title":"Ninth European Conference on Speech Communication and Technology.","author":"Black A. W.","unstructured":"A. W. Black and K. Tokuda . 2005. The Blizzard Challenge-2005: Evaluating corpus-based speech synthesis on common datasets . In Ninth European Conference on Speech Communication and Technology. A. W. Black and K. Tokuda. 2005. The Blizzard Challenge-2005: Evaluating corpus-based speech synthesis on common datasets. In Ninth European Conference on Speech Communication and Technology."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-1333"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"crossref","unstructured":"D. A. Braude M. P. Aylett C. Laoide-Kemp S. Ashby K. M. Scott B. O. Raghallaigh A. Braudo A. Brouwer and A. Stan. 2019. All together now: The living audio dataset. In Interspeech. 1521\u20131525. D. A. Braude M. P. Aylett C. Laoide-Kemp S. Ashby K. M. Scott B. O. Raghallaigh A. Braudo A. Brouwer and A. Stan. 2019. All together now: The living audio dataset. In Interspeech . 1521\u20131525.","DOI":"10.21437\/Interspeech.2019-2448"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/Humanoids.2011.6100810"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijhcs.2015.01.006"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3359325"},{"key":"e_1_3_2_1_27_1","volume-title":"Working with Spoken Discourse","author":"Cameron D.","unstructured":"D. Cameron . 2001. Working with Spoken Discourse . Sage . D. Cameron. 2001. Working with Spoken Discourse. Sage."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.tics.2007.10.001"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0271-5309(97)00016-5"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2983926"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1093\/iwc\/iwz016"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300705"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"crossref","unstructured":"M. Cooke C. Mayo and C. Valentini-Botinhao. 2013 August. Intelligibility-enhancing speech modifications: The Hurricane Challenge. In Interspeech. 3552\u20133556. M. Cooke C. Mayo and C. Valentini-Botinhao. 2013 August. Intelligibility-enhancing speech modifications: The Hurricane Challenge. In Interspeech . 3552\u20133556.","DOI":"10.21437\/Interspeech.2013-764"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2935334.2935386"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-9841.2007.00311.x"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3098279.3098539"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3342775.3342786"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9781139166973"},{"key":"e_1_3_2_1_39_1","volume-title":"A Dictionary of Linguistics and Phonetics","author":"Crystal D.","unstructured":"D. Crystal . 1997. A Dictionary of Linguistics and Phonetics . Blackwell , UK. D. Crystal. 1997. A Dictionary of Linguistics and Phonetics. Blackwell, UK."},{"key":"e_1_3_2_1_40_1","volume-title":"A Dictionary of Linguistics and Phonetics","author":"Crystal D.","unstructured":"D. Crystal . 2011. A Dictionary of Linguistics and Phonetics , Vol. 30 . John Wiley & Sons . D. Crystal. 2011. A Dictionary of Linguistics and Phonetics, Vol. 30. John Wiley & Sons."},{"key":"e_1_3_2_1_41_1","volume-title":"Proceedings of Interact. 294\u2013301","author":"Dahlb\u00e4ck N.","unstructured":"N. Dahlb\u00e4ck , S. Swamy , C. Nass , F. Arvidsson , and J. Sk\u00e5geby . 2001. Spoken interaction with computers in a native or non-native language\u2014Same or different . In Proceedings of Interact. 294\u2013301 . N. Dahlb\u00e4ck, S. Swamy, C. Nass, F. Arvidsson, and J. Sk\u00e5geby. 2001. Spoken interaction with computers in a native or non-native language\u2014Same or different. In Proceedings of Interact. 294\u2013301."},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/1240624.1240859"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3405755.3406151"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.specom.2013.10.002"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3319502.3374815"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.5555\/2615731.2617415"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3338286.3340116"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445206"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3342775.3342785"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISSPIT.2015.7394422"},{"key":"e_1_3_2_1_51_1","volume-title":"2015 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG)","volume":"5","author":"Haddad K. El","unstructured":"K. El Haddad , S. Dupont , N. D\u2019Alessandro , and T. Dutoit . 2015b. An HMM-based speech\u2013smile synthesis system: An approach for amusement synthesis . In 2015 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG) , Vol. 5 . IEEE, 1\u20136. K. El Haddad, S. Dupont, N. D\u2019Alessandro, and T. Dutoit. 2015b. An HMM-based speech\u2013smile synthesis system: An approach for amusement synthesis. In 2015 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG), Vol. 5. IEEE, 1\u20136."},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-68456-7"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1250\/ast.26.317"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0921-8890(02)00372-X"},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.5555\/2927507"},{"key":"e_1_3_2_1_56_1","unstructured":"K. Georgila A. W. Black K. Sagae and D. R. Traum. 2012. Practical evaluation of human and synthesized speech for virtual human dialogue systems. In LREC. 3519\u20133526. K. Georgila A. W. Black K. Sagae and D. R. Traum. 2012. Practical evaluation of human and synthesized speech for virtual human dialogue systems. In LREC . 3519\u20133526."},{"key":"e_1_3_2_1_57_1","volume-title":"Interaction Ritual: Essays in Face-to Face-Behavior. AldineTransaction.","author":"Goffman E.","year":"2005","unstructured":"E. Goffman . 2005 . Interaction Ritual: Essays in Face-to Face-Behavior. AldineTransaction. E. Goffman. 2005. Interaction Ritual: Essays in Face-to Face-Behavior. AldineTransaction."},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2018-1174"},{"key":"e_1_3_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10772-012-9180-2"},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.5555\/3141475.3141498"},{"key":"e_1_3_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.21437\/SSW.2019-43"},{"key":"e_1_3_2_1_62_1","volume-title":"Proc. Interspeech.","author":"Hofer G.","unstructured":"G. Hofer , K. Richmond , and R. Clark . 2005. Informed blending of databases for emotional speech synthesis . In Proc. Interspeech. G. Hofer, K. Richmond, and R. Clark. 2005. Informed blending of databases for emotional speech synthesis. In Proc. Interspeech."},{"key":"e_1_3_2_1_63_1","doi-asserted-by":"crossref","unstructured":"A. Hughes P. Trudgill and D. Watt. 2013. English Accents and Dialects: An Introduction to Social and Regional Varieties of English in the British Isles. Routledge. A. Hughes P. Trudgill and D. Watt. 2013. English Accents and Dialects: An Introduction to Social and Regional Varieties of English in the British Isles . Routledge.","DOI":"10.4324\/9780203784440"},{"key":"e_1_3_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1996.541110"},{"key":"e_1_3_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1155\/2007\/76030"},{"key":"e_1_3_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3313831.3376863"},{"key":"e_1_3_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.3758\/s13423-019-01701-x"},{"key":"e_1_3_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-13-9443-0_6"},{"key":"e_1_3_2_1_69_1","volume-title":"International Conference on Machine Learning. 3331\u20133340","author":"Kenter T.","unstructured":"T. Kenter , V. Wan , C.-A. Chan , R. Clark , and J. Vit . 2019. CHiVE: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network . In International Conference on Machine Learning. 3331\u20133340 . T. Kenter, V. Wan, C.-A. Chan, R. Clark, and J. Vit. 2019. CHiVE: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network. In International Conference on Machine Learning. 3331\u20133340."},{"key":"e_1_3_2_1_70_1","volume-title":"2004 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), (IEEE Cat. No. 04CH37566)","volume":"4","author":"Kidd C. D.","unstructured":"C. D. Kidd and C. Breazeal . 2004. Effect of a robot on user perceptions . In 2004 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), (IEEE Cat. No. 04CH37566) . Vol. 4 . IEEE, 3559\u20133564. C. D. Kidd and C. Breazeal. 2004. Effect of a robot on user perceptions. In 2004 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), (IEEE Cat. No. 04CH37566). Vol. 4. IEEE, 3559\u20133564."},{"key":"e_1_3_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-7687.2010.00965.x"},{"key":"e_1_3_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/3405755.3406119"},{"key":"e_1_3_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.383940"},{"key":"e_1_3_2_1_74_1","volume-title":"Fifth ISCA Workshop on Speech Synthesis.","author":"Kominek J.","unstructured":"J. Kominek and A. W. Black . 2004. The CMU Arctic speech databases . In Fifth ISCA Workshop on Speech Synthesis. J. Kominek and A. W. Black. 2004. The CMU Arctic speech databases. In Fifth ISCA Workshop on Speech Synthesis."},{"key":"e_1_3_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10772-011-9125-1"},{"key":"e_1_3_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1163\/016918609X12518783330360"},{"key":"e_1_3_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jml.2007.06.005"},{"key":"e_1_3_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.apergo.2017.04.003"},{"key":"e_1_3_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/2470654.2466455"},{"key":"e_1_3_2_1_80_1","volume-title":"Sixth International Conference on Spoken Language Processing.","author":"Lenzo K. A.","unstructured":"K. A. Lenzo and A. W. Black . 2000. Diphone collection and synthesis . In Sixth International Conference on Spoken Language Processing. K. A. Lenzo and A. W. Black. 2000. Diphone collection and synthesis. In Sixth International Conference on Spoken Language Processing."},{"key":"e_1_3_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijhcs.2015.01.001"},{"key":"e_1_3_2_1_82_1","doi-asserted-by":"crossref","unstructured":"J. Lorenzo-Trueba T. Drugman J. Latorre T. Merritt B. Putrycz R. Barra-Chicote Alexis Moinet and Vatsal Aggarwal. 2018. Towards achieving robust universal neural vocoding. arXiv preprint arXiv:1811.06292. J. Lorenzo-Trueba T. Drugman J. Latorre T. Merritt B. Putrycz R. Barra-Chicote Alexis Moinet and Vatsal Aggarwal. 2018. Towards achieving robust universal neural vocoding. arXiv preprint arXiv:1811.06292 .","DOI":"10.21437\/Interspeech.2019-1424"},{"key":"e_1_3_2_1_83_1","doi-asserted-by":"publisher","DOI":"10.1145\/2858036.2858288"},{"key":"e_1_3_2_1_84_1","volume-title":"Nautilus: A versatile voice cloning system. arXiv preprint arXiv:2005.11004.","author":"Luong H.-T.","year":"2020","unstructured":"H.-T. Luong and J. Yamagishi . 2020 . Nautilus: A versatile voice cloning system. arXiv preprint arXiv:2005.11004. H.-T. Luong and J. Yamagishi. 2020. Nautilus: A versatile voice cloning system. arXiv preprint arXiv:2005.11004."},{"key":"e_1_3_2_1_85_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683815"},{"key":"e_1_3_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.5555\/3378680.3378713"},{"key":"e_1_3_2_1_87_1","doi-asserted-by":"publisher","DOI":"10.1038\/264746a0"},{"key":"e_1_3_2_1_88_1","doi-asserted-by":"publisher","DOI":"10.1145\/1518701.1518970"},{"key":"e_1_3_2_1_89_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-1438"},{"key":"e_1_3_2_1_90_1","volume-title":"Dialogues with Social Robots","author":"Moore R. K.","unstructured":"R. K. Moore . 2017. Is spoken language all-or-nothing? Implications for future speech-based human\u2013machine interaction . In Dialogues with Social Robots . Springer , 281\u2013291. R. K. Moore. 2017. Is spoken language all-or-nothing? Implications for future speech-based human\u2013machine interaction. In Dialogues with Social Robots. Springer, 281\u2013291."},{"key":"e_1_3_2_1_91_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2016-874"},{"key":"e_1_3_2_1_92_1","doi-asserted-by":"publisher","DOI":"10.1145\/1957656.1957786"},{"key":"e_1_3_2_1_93_1","doi-asserted-by":"publisher","DOI":"10.1145\/332040.332452"},{"key":"e_1_3_2_1_94_1","doi-asserted-by":"publisher","DOI":"10.1037\/\/1076-898X.7.3.171"},{"key":"e_1_3_2_1_95_1","doi-asserted-by":"publisher","DOI":"10.5555\/1088925"},{"key":"e_1_3_2_1_96_1","doi-asserted-by":"publisher","DOI":"10.1145\/191666.191703"},{"key":"e_1_3_2_1_97_1","doi-asserted-by":"publisher","DOI":"10.1145\/1463160.1463235"},{"key":"e_1_3_2_1_98_1","doi-asserted-by":"publisher","DOI":"10.1159\/000261678"},{"key":"e_1_3_2_1_99_1","unstructured":"A. V. D. Oord S. Dieleman H. Zen K. Simonyan O. Vinyals A. Graves N. Kalchbrenner A. Senior and K. Kavukcuoglu. 2016. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499. A. V. D. Oord S. Dieleman H. Zen K. Simonyan O. Vinyals A. Graves N. Kalchbrenner A. Senior and K. Kavukcuoglu. 2016. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 ."},{"key":"e_1_3_2_1_100_1","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511571299"},{"key":"e_1_3_2_1_101_1","doi-asserted-by":"publisher","DOI":"10.1017\/S0047404505050037"},{"key":"e_1_3_2_1_102_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W15-4613"},{"key":"e_1_3_2_1_103_1","unstructured":"W. Ping K. Peng A. Gibiansky S. O. Arik A. Kannan S. Narang J. Raiman and J. Miller. 2017. Deep Voice 3: Scaling text-to-speech with convolutional sequence learning. arXiv preprint arXiv:1710.07654. W. Ping K. Peng A. Gibiansky S. O. Arik A. Kannan S. Narang J. Raiman and J. Miller. 2017. Deep Voice 3: Scaling text-to-speech with convolutional sequence learning. arXiv preprint arXiv:1710.07654 ."},{"key":"e_1_3_2_1_104_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2006.876123"},{"key":"e_1_3_2_1_105_1","doi-asserted-by":"publisher","DOI":"10.1145\/2998181.2998298"},{"key":"e_1_3_2_1_106_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173574.3174214"},{"key":"e_1_3_2_1_107_1","volume-title":"International Conference on Intelligent Virtual Agents. Springer, 190\u2013197","author":"Potard B.","unstructured":"B. Potard , M. P. Aylett , and D. A. Braude . 2016. Cross modal evaluation of high quality emotional speech synthesis with the Virtual Human Toolkit . In International Conference on Intelligent Virtual Agents. Springer, 190\u2013197 . B. Potard, M. P. Aylett, and D. A. Braude. 2016. Cross modal evaluation of high quality emotional speech synthesis with the Virtual Human Toolkit. In International Conference on Intelligent Virtual Agents. Springer, 190\u2013197."},{"key":"e_1_3_2_1_108_1","doi-asserted-by":"crossref","unstructured":"N. Prateek M. \u0141ajszczak R. Barra-Chicote T. Drugman J. Lorenzo-Trueba T. Merritt S. Ronanki and T. Wood. 2019. In other news: A bi-style text-to-speech model for synthesizing newscaster voice with limited data. arXiv preprint arXiv:1904.02790. N. Prateek M. \u0141ajszczak R. Barra-Chicote T. Drugman J. Lorenzo-Trueba T. Merritt S. Ronanki and T. Wood. 2019. In other news: A bi-style text-to-speech model for synthesizing newscaster voice with limited data. arXiv preprint arXiv:1904.02790 .","DOI":"10.18653\/v1\/N19-2026"},{"key":"e_1_3_2_1_109_1","doi-asserted-by":"crossref","unstructured":"S. Ramakrishnan. 2012. Recognition of emotion from speech: A review. Speech Enhancement Modeling and Recognition\u2013Algorithms and Applications 7 121\u2013137. S. Ramakrishnan. 2012. Recognition of emotion from speech: A review. Speech Enhancement Modeling and Recognition\u2013Algorithms and Applications 7 121\u2013137.","DOI":"10.5772\/39246"},{"key":"e_1_3_2_1_110_1","doi-asserted-by":"publisher","DOI":"10.1145\/3313831.3376767"},{"key":"e_1_3_2_1_111_1","doi-asserted-by":"publisher","DOI":"10.1080\/09658210903130764"},{"key":"e_1_3_2_1_112_1","unstructured":"E. B. Ryan and H. Giles. 1982. An integrative perspective for the study of attitudes towards language variation. In Attitudes Towards Language Variation: Social and Applied Contexts. Edward Arnold London 1\u201319. E. B. Ryan and H. Giles. 1982. An integrative perspective for the study of attitudes towards language variation. In Attitudes Towards Language Variation: Social and Applied Contexts . Edward Arnold London 1\u201319."},{"key":"e_1_3_2_1_113_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12369-013-0196-9"},{"key":"e_1_3_2_1_114_1","doi-asserted-by":"publisher","DOI":"10.1145\/1978942.1979353"},{"key":"e_1_3_2_1_115_1","doi-asserted-by":"publisher","DOI":"10.1145\/3342775.3342803"},{"key":"e_1_3_2_1_116_1","volume-title":"Proceedings Eurospeech 01","author":"Schr\u00f6der M.","year":"2001","unstructured":"M. Schr\u00f6der . 2001 . Emotional speech synthesis: A review . In Proceedings Eurospeech 01 . 561\u2013564. M. Schr\u00f6der. 2001. Emotional speech synthesis: A review. In Proceedings Eurospeech 01. 561\u2013564."},{"key":"e_1_3_2_1_117_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-24842-2_21"},{"key":"e_1_3_2_1_118_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-84800-306-4_7"},{"key":"e_1_3_2_1_119_1","doi-asserted-by":"publisher","DOI":"10.1109\/T-AFFC.2011.34"},{"key":"e_1_3_2_1_120_1","unstructured":"R. Skerry-Ryan E. Battenberg Y. Xiao Y. Wang D. Stanton J. Shor R. J. Weiss R. Clark and R. A. Saurous. 2018. Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron. arXiv preprint arXiv:1803.09047. R. Skerry-Ryan E. Battenberg Y. Xiao Y. Wang D. Stanton J. Shor R. J. Weiss R. Clark and R. A. Saurous. 2018. Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron. arXiv preprint arXiv:1803.09047 ."},{"key":"e_1_3_2_1_121_1","doi-asserted-by":"publisher","DOI":"10.1145\/1957656.1957757"},{"key":"e_1_3_2_1_122_1","volume-title":"ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 6264\u20136268","author":"Sun G.","unstructured":"G. Sun , Y. Zhang , R. J. Weiss , Y. Cao , H. Zen , and Y. Wu . 2020. Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis . In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 6264\u20136268 . G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, and Y. Wu. 2020. Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 6264\u20136268."},{"key":"e_1_3_2_1_123_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.2390679"},{"key":"e_1_3_2_1_124_1","doi-asserted-by":"publisher","DOI":"10.1145\/3405755.3406123"},{"key":"e_1_3_2_1_125_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.wocn.2009.10.002"},{"key":"e_1_3_2_1_127_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-2836"},{"key":"e_1_3_2_1_128_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.410151"},{"key":"e_1_3_2_1_129_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-6393(96)00068-4"},{"key":"e_1_3_2_1_130_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2006.876129"},{"key":"e_1_3_2_1_131_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2000.861820"},{"key":"e_1_3_2_1_132_1","doi-asserted-by":"publisher","DOI":"10.1109\/RO-MAN47096.2020.9223599"},{"key":"e_1_3_2_1_133_1","volume-title":"Proceedings of the 18th International Congress of Phonetic Sciences (ICPhS","author":"Torre I.","year":"2015","unstructured":"I. Torre , J. Goslin , and L. White . 2015. Investing in accents: How does experience mediate trust attributions to different voices? In Proceedings of the 18th International Congress of Phonetic Sciences (ICPhS 2015 ). I. Torre, J. Goslin, and L. White. 2015. Investing in accents: How does experience mediate trust attributions to different voices? In Proceedings of the 18th International Congress of Phonetic Sciences (ICPhS 2015)."},{"key":"e_1_3_2_1_134_1","doi-asserted-by":"publisher","DOI":"10.1145\/3242969.3242984"},{"key":"e_1_3_2_1_135_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.chb.2019.106215"},{"key":"e_1_3_2_1_136_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-24842-2_23"},{"key":"e_1_3_2_1_137_1","unstructured":"G. R. Tucker. 1999. A Global Perspective on Bilingualism and Bilingual Education. ERIC Digest. G. R. Tucker. 1999. A Global Perspective on Bilingualism and Bilingual Education . ERIC Digest."},{"key":"e_1_3_2_1_138_1","doi-asserted-by":"publisher","DOI":"10.21437\/SSW.2019-19"},{"key":"e_1_3_2_1_139_1","doi-asserted-by":"publisher","DOI":"10.1109\/ROMAN.2006.314404"},{"key":"e_1_3_2_1_140_1","volume-title":"Tacotron: Towards end-to-end speech synthesis. arXiv preprint arXiv:1703.10135.","author":"Wang Y.","year":"2017","unstructured":"Y. Wang , R. Skerry-Ryan , D. Stanton , Y. Wu , R. J. Weiss , N. Jaitly , Z. Yang , Y. Xiao , Z. Chen , S. Bengio , Q. Le , Y. Agiomyrgiannakis , R. Clark , and R. A. Saurou . 2017 . Tacotron: Towards end-to-end speech synthesis. arXiv preprint arXiv:1703.10135. Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurou. 2017. Tacotron: Towards end-to-end speech synthesis. arXiv preprint arXiv:1703.10135."},{"key":"e_1_3_2_1_141_1","unstructured":"M. West R. Kraut and H. Chew 2019. I\u2019d blush if I could: Closing gender divides in digital skills through education. https:\/\/unesdoc.unesco.org\/ark:\/48223\/pf0000367416. M. West R. Kraut and H. Chew 2019. I\u2019d blush if I could: Closing gender divides in digital skills through education. https:\/\/unesdoc.unesco.org\/ark:\/48223\/pf0000367416."},{"key":"e_1_3_2_1_142_1","doi-asserted-by":"crossref","unstructured":"M. Wester M. Aylett M. Tomalin and R. Dall. 2015. Artificial personality and disfluency. Interspeech 2015. 5. M. Wester M. Aylett M. Tomalin and R. Dall. 2015. Artificial personality and disfluency. Interspeech 2015 . 5.","DOI":"10.21437\/Interspeech.2015-141"},{"key":"e_1_3_2_1_143_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-1250"},{"key":"e_1_3_2_1_144_1","doi-asserted-by":"publisher","DOI":"10.1145\/3405755.3406154"},{"key":"e_1_3_2_1_145_1","volume-title":"Blizzard Challenge Workshop. http:\/\/www.festvox.org\/blizzard\/bc2019\/blizzard2019_overview_paper.pdf.","author":"Wu Z.","unstructured":"Z. Wu , Z. Xie , and S. King . 2019. The blizzard challenge 2019 . Blizzard Challenge Workshop. http:\/\/www.festvox.org\/blizzard\/bc2019\/blizzard2019_overview_paper.pdf. Z. Wu, Z. Xie, and S. King. 2019. The blizzard challenge 2019. Blizzard Challenge Workshop. http:\/\/www.festvox.org\/blizzard\/bc2019\/blizzard2019_overview_paper.pdf."},{"key":"e_1_3_2_1_146_1","doi-asserted-by":"publisher","DOI":"10.1145\/3405755.3406118"},{"key":"e_1_3_2_1_147_1","doi-asserted-by":"publisher","DOI":"10.1145\/3379503.3403563"},{"key":"e_1_3_2_1_148_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2006.1659961"},{"key":"e_1_3_2_1_149_1","doi-asserted-by":"publisher","DOI":"10.7488\/ds\/2645"},{"key":"e_1_3_2_1_150_1","volume-title":"2013 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 7962\u20137966","author":"Zen H.","unstructured":"H. Zen , A. Senior , and M. Schuster . 2013. Statistical parametric speech synthesis using deep neural networks . In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 7962\u20137966 . H. Zen, A. Senior, and M. Schuster. 2013. Statistical parametric speech synthesis using deep neural networks. In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE, 7962\u20137966."},{"key":"e_1_3_2_1_151_1","doi-asserted-by":"crossref","unstructured":"H. Zen V. Dang R. Clark Y. Zhang R. J. Weiss Y. Jia Z. Chen and Y. Wu. 2019. LibriTTS: A corpus derived from LibriSpeech for text-to-speech. arXiv preprint arXiv:1904.02882. H. Zen V. Dang R. Clark Y. Zhang R. J. Weiss Y. Jia Z. Chen and Y. Wu. 2019. LibriTTS: A corpus derived from LibriSpeech for text-to-speech. arXiv preprint arXiv:1904.02882 .","DOI":"10.21437\/Interspeech.2019-2441"},{"key":"e_1_3_2_1_152_1","doi-asserted-by":"crossref","unstructured":"Y. Zhang R. J. Weiss H. Zen Y. Wu Z. Chen R. Skerry-Ryan Y. Jia A. Rosenberg and B. Ramabhadran. 2019. Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning. arXiv preprint arXiv:1907.04448. Y. Zhang R. J. Weiss H. Zen Y. Wu Z. Chen R. Skerry-Ryan Y. Jia A. Rosenberg and B. Ramabhadran. 2019. Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning. arXiv preprint arXiv:1907.04448 .","DOI":"10.21437\/Interspeech.2019-2668"},{"key":"e_1_3_2_1_153_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cogsys.2019.09.009"}],"container-title":["The Handbook on Socially Interactive Agents"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3477322.3477329","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3477322.3477329","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:30:31Z","timestamp":1750188631000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3477322.3477329"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,9,10]]},"ISBN":["9781450387200"],"references-count":152,"alternative-id":["10.1145\/3477322.3477329","10.1145\/3477322"],"URL":"https:\/\/doi.org\/10.1145\/3477322.3477329","relation":{},"subject":[],"published":{"date-parts":[[2021,9,10]]},"assertion":[{"value":"2021-10-02","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-10-02","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}