{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T10:10:40Z","timestamp":1781086240140,"version":"3.54.1"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2017,7,20]],"date-time":"2017-07-20T00:00:00Z","timestamp":1500508800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100000266","name":"EPSRC","doi-asserted-by":"crossref","award":["EP\/M014053\/1"],"award-info":[{"award-number":["EP\/M014053\/1"]}],"id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2017,8,31]]},"abstract":"<jats:p>We introduce a simple and effective deep learning approach to automatically generate natural looking speech animation that synchronizes to input speech. Our approach uses a sliding window predictor that learns arbitrary nonlinear mappings from phoneme label input sequences to mouth movements in a way that accurately captures natural motion and visual coarticulation effects. Our deep learning approach enjoys several attractive properties: it runs in real-time, requires minimal parameter tuning, generalizes well to novel input speech sequences, is easily edited to create stylized and emotional speech, and is compatible with existing animation retargeting approaches. One important focus of our work is to develop an effective approach for speech animation that can be easily integrated into existing production pipelines. We provide a detailed description of our end-to-end approach, including machine learning design decisions. Generalized speech animation results are demonstrated over a wide range of animation clips on a variety of characters and voices, including singing and foreign language input. Our approach can also generate on-demand speech animation in real-time from user speech input.<\/jats:p>","DOI":"10.1145\/3072959.3073699","type":"journal-article","created":{"date-parts":[[2017,7,21]],"date-time":"2017-07-21T12:24:07Z","timestamp":1500639847000},"page":"1-11","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":230,"title":["A deep learning approach for generalized speech animation"],"prefix":"10.1145","volume":"36","author":[{"given":"Sarah","family":"Taylor","sequence":"first","affiliation":[{"name":"University of East Anglia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Taehwan","family":"Kim","sequence":"additional","affiliation":[{"name":"California Institute of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yisong","family":"Yue","sequence":"additional","affiliation":[{"name":"California Institute of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Moshe","family":"Mahler","sequence":"additional","affiliation":[{"name":"Disney Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"James","family":"Krahe","sequence":"additional","affiliation":[{"name":"Disney Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anastasio Garcia","family":"Rodriguez","sequence":"additional","affiliation":[{"name":"Disney Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jessica","family":"Hodgins","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Iain","family":"Matthews","sequence":"additional","affiliation":[{"name":"Disney Research"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2017,7,20]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.434"},{"key":"e_1_2_2_2_1","volume-title":"Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473","author":"Bahdanau Dzmitry","year":"2014","unstructured":"Dzmitry Bahdanau , Kyunghyun Cho , and Yoshua Bengio . 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473 ( 2014 ). Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473 (2014)."},{"key":"e_1_2_2_3_1","volume-title":"Deep Learning and Unsupervised Feature Learning NIPS 2012 Workshop.","author":"Bastien Fr\u00e9d\u00e9ric","year":"2012","unstructured":"Fr\u00e9d\u00e9ric Bastien , Pascal Lamblin , Razvan Pascanu , James Bergstra , Ian Goodfellow , Arnaud Bergeron , Nicolas Bouchard , David Warde-Farley , and Yoshua Bengio . 2012 . Theano: new features and speed improvements . Deep Learning and Unsupervised Feature Learning NIPS 2012 Workshop. (2012). Fr\u00e9d\u00e9ric Bastien, Pascal Lamblin, Razvan Pascanu, James Bergstra, Ian Goodfellow, Arnaud Bergeron, Nicolas Bouchard, David Warde-Farley, and Yoshua Bengio. 2012. Theano: new features and speed improvements. Deep Learning and Unsupervised Feature Learning NIPS 2012 Workshop. (2012)."},{"key":"e_1_2_2_4_1","volume-title":"High-quality passive facial performance capture using anchor frames. ACM Transactions on Graphics 30 (Aug","author":"Beeler Thabo","year":"2011","unstructured":"Thabo Beeler , Fabian Hahn , Derek Bradley , Bernd Bickel , Paul Beardsley , Craig Gotsman , Robert W Sumner , and Markus Gross . 2011. High-quality passive facial performance capture using anchor frames. ACM Transactions on Graphics 30 (Aug . 2011 ), 75:1--75:10. Issue 4. Thabo Beeler, Fabian Hahn, Derek Bradley, Bernd Bickel, Paul Beardsley, Craig Gotsman, Robert W Sumner, and Markus Gross. 2011. High-quality passive facial performance capture using anchor frames. ACM Transactions on Graphics 30 (Aug. 2011), 75:1--75:10. Issue 4."},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/311535.311537"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/258734.258880"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766943"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2462012"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1095878.1095881"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143865"},{"key":"e_1_2_2_11_1","volume-title":"Models and Techniques in Computer Animation","author":"Cohen Michael M","unstructured":"Michael M Cohen , Dominic W Massaro , and others. 1994. Modeling Coarticualtion in Synthetic Visual Speech . In Models and Techniques in Computer Animation , N.M. Thalmann and Thalmann D (Eds.). Springer-Verlag , 141--155. Michael M Cohen, Dominic W Massaro, and others. 1994. Modeling Coarticualtion in Synthetic Visual Speech. In Models and Techniques in Computer Animation, N.M. Thalmann and Thalmann D (Eds.). Springer-Verlag, 141--155."},{"key":"e_1_2_2_12_1","first-page":"2493","article-title":"Natural language processing (almost) from scratch","author":"Collobert Ronan","year":"2011","unstructured":"Ronan Collobert , Jason Weston , L\u00e9on Bottou , Michael Karlen , Koray Kavukcuoglu , and Pavel Kuksa . 2011 . Natural language processing (almost) from scratch . Journal of Machine Learning Research 12 , Aug (2011), 2493 -- 2537 . Ronan Collobert, Jason Weston, L\u00e9on Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011. Natural language processing (almost) from scratch. Journal of Machine Learning Research 12, Aug (2011), 2493--2537.","journal-title":"Journal of Machine Learning Research 12"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.927467"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/6046.865480"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cag.2006.08.017"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1891903.1891942"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925984"},{"key":"e_1_2_2_18_1","volume-title":"Proceedings of Advances in Natural Information Processing Systems. 401--408","author":"Englebienne Gwenn","year":"2007","unstructured":"Gwenn Englebienne , Timothy F Cootes , and Magnus Rattray . 2007 . A Probabilistic Model for Generating Realistic Speech Movements from Speech . In Proceedings of Advances in Natural Information Processing Systems. 401--408 . Gwenn Englebienne, Timothy F Cootes, and Magnus Rattray. 2007. A Probabilistic Model for Generating Realistic Speech Movements from Speech. In Proceedings of Advances in Natural Information Processing Systems. 401--408."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/566570.566594"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178899"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2005.843341"},{"key":"e_1_2_2_22_1","first-page":"8","article-title":"Driving High-Resolution Facial Scans with Video Performance Capture","volume":"34","author":"Fyfe Graham","year":"2014","unstructured":"Graham Fyfe , Andrew Jones , Oleg Alexander , Ryosuke Ichikari , and Paul Debevec . 2014 . Driving High-Resolution Facial Scans with Video Performance Capture . ACM Transactions on Graphics 34 , 1 (2014), 8 . Graham Fyfe, Andrew Jones, Oleg Alexander, Ryosuke Ichikari, and Paul Debevec. 2014. Driving High-Resolution Facial Scans with Video Performance Capture. ACM Transactions on Graphics 34, 1 (2014), 8.","journal-title":"ACM Transactions on Graphics"},{"key":"e_1_2_2_24_1","volume-title":"Proceedings of Interspeech. 2474--2477","author":"Govokhina Oxana","year":"2006","unstructured":"Oxana Govokhina , G\u00e9rard Bailly , Gaspard Breton , and Paul Bagshaw . 2006 . TDA: A new trainable trajectory formation system for facial animation . In Proceedings of Interspeech. 2474--2477 . Oxana Govokhina, G\u00e9rard Bailly, Gaspard Breton, and Paul Bagshaw. 2006. TDA: A new trainable trajectory formation system for facial animation. In Proceedings of Interspeech. 2474--2477."},{"key":"e_1_2_2_25_1","first-page":"1764","article-title":"Towards End-To-End Speech Recognition with Recurrent Neural Networks","volume":"14","author":"Graves Alex","year":"2014","unstructured":"Alex Graves and Navdeep Jaitly . 2014 . Towards End-To-End Speech Recognition with Recurrent Neural Networks . In ICML , Vol. 14. 1764 -- 1772 . Alex Graves and Navdeep Jaitly. 2014. Towards End-To-End Speech Recognition with Recurrent Neural Networks. In ICML, Vol. 14. 1764--1772.","journal-title":"ICML"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/1964921.1964969"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2783258.2783356"},{"key":"e_1_2_2_29_1","unstructured":"Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Neural Information Processing Systems. 1097--1105.  Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Neural Information Processing Systems. 1097--1105."},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2462019"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/s12193-011-0070-8"},{"key":"e_1_2_2_32_1","volume-title":"IEEE Conference on Multimedia and Expo Workshops. 1--6.","author":"Luo Changwei","year":"2014","unstructured":"Changwei Luo , Jun Yu , Xian Li , and Zengfu Wang . 2014 . Realtime speech-driven facial animation using Gaussian Mixture Models . In IEEE Conference on Multimedia and Expo Workshops. 1--6. Changwei Luo, Jun Yu, Xian Li, and Zengfu Wang. 2014. Realtime speech-driven facial animation using Gaussian Mixture Models. In IEEE Conference on Multimedia and Expo Workshops. 1--6."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2006.18"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000029666.37597.d3"},{"key":"e_1_2_2_35_1","first-page":"7","article-title":"Comprehensive many-to-many phoneme-to-viseme mapping and its application for concatenative visual speech synthesis","volume":"55","author":"Mattheyses Wesley","year":"2013","unstructured":"Wesley Mattheyses , Lukas Latacz , and Werner Verhelst . 2013 . Comprehensive many-to-many phoneme-to-viseme mapping and its application for concatenative visual speech synthesis . Speech Communication 55 , 7 -- 8 (2013), 857--876. Wesley Mattheyses, Lukas Latacz, and Werner Verhelst. 2013. Comprehensive many-to-many phoneme-to-viseme mapping and its application for concatenative visual speech synthesis. Speech Communication 55, 7--8 (2013), 857--876.","journal-title":"Speech Communication"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1038\/264746a0"},{"key":"e_1_2_2_37_1","volume-title":"ISCA Workshop on Speech Synthesis. 185--190","author":"Merritt Thomas","year":"2013","unstructured":"Thomas Merritt and Simon King . 2013 . Investigating the shortcomings of HMM synthesis . In ISCA Workshop on Speech Synthesis. 185--190 . Thomas Merritt and Simon King. 2013. Investigating the shortcomings of HMM synthesis. In ISCA Workshop on Speech Synthesis. 185--190."},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2037715.2037724"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2670313"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015706.1015736"},{"key":"e_1_2_2_42_1","unstructured":"Ilya Sutskever Oriol Vinyals and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In Neural Information Processing Systemsw. 3104--3112.  Ilya Sutskever Oriol Vinyals and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In Neural Information Processing Systemsw. 3104--3112."},{"key":"e_1_2_2_43_1","volume-title":"Proceedings of ACM SIGGRAPH\/Eurographics Symposium on Computer Animation. Eurographics Association, 275--284","author":"Taylor Sarah L","year":"2012","unstructured":"Sarah L Taylor , Moshe Mahler , Barry-John Theobald , and Iain Matthews . 2012 . Dynamic Units of Visual Speech . In Proceedings of ACM SIGGRAPH\/Eurographics Symposium on Computer Animation. Eurographics Association, 275--284 . Sarah L Taylor, Moshe Mahler, Barry-John Theobald, and Iain Matthews. 2012. Dynamic Units of Visual Speech. In Proceedings of ACM SIGGRAPH\/Eurographics Symposium on Computer Animation. Eurographics Association, 275--284."},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2012.2202651"},{"key":"e_1_2_2_45_1","volume-title":"Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499","author":"van den Oord A\u00e4ron","year":"2016","unstructured":"A\u00e4ron van den Oord , Sander Dieleman , Heiga Zen , Karen Simonyan , Oriol Vinyals , Alex Graves , Nal Kalchbrenner , Andrew Senior , and Koray Kavukcuoglu . 2016 . Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 (2016). A\u00e4ron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 (2016)."},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2012.6288925"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/1964921.1964972"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.gmod.2013.10.002"},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2006.12.001"},{"key":"e_1_2_2_50_1","volume-title":"attend and tell: Neural image caption generation with visual attention. arXiv preprint arXiv:1502.03044 2, 3","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu , Jimmy Ba , Ryan Kiros , Kyunghyun Cho , Aaron Courville , Ruslan Salakhudinov , Rich Zemel , and Yoshua Bengio . 2015. Show , attend and tell: Neural image caption generation with visual attention. arXiv preprint arXiv:1502.03044 2, 3 ( 2015 ), 5. Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. arXiv preprint arXiv:1502.03044 2, 3 (2015), 5."},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2522628.2522904"},{"key":"e_1_2_2_52_1","volume-title":"The HTK Book","author":"Young Steve","unstructured":"Steve Young , Gunnar Evermann , Mark Gales , Thomas Hain , Dan Kershaw , Xunying Liu , Gareth Moore , Julian Odell , Dave Ollason , Dan Povey , and others. 2006. The HTK Book . Cambridge University . Steve Young, Gunnar Evermann, Mark Gales, Thomas Hain, Dan Kershaw, Xunying Liu, Gareth Moore, Julian Odell, Dave Ollason, Dan Povey, and others. 2006. The HTK Book. Cambridge University."},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.2935783"},{"key":"e_1_2_2_54_1","volume-title":"Proceedings of the Speech Synthesis Workshop. 294--299","author":"Zen Heiga","year":"2007","unstructured":"Heiga Zen , Takashi Nose , Junichi Yamagishi , Shinji Sako , Takashi Masuko , Alan Black , and Keiichi Tokuda . 2007 . The HMM-based speech synthesis system version 2.0 . In Proceedings of the Speech Synthesis Workshop. 294--299 . Heiga Zen, Takashi Nose, Junichi Yamagishi, Shinji Sako, Takashi Masuko, Alan Black, and Keiichi Tokuda. 2007. The HMM-based speech synthesis system version 2.0. In Proceedings of the Speech Synthesis Workshop. 294--299."},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/1186562.1015759"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3072959.3073699","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3072959.3073699","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:30:23Z","timestamp":1750217423000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3072959.3073699"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,7,20]]},"references-count":53,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2017,8,31]]}},"alternative-id":["10.1145\/3072959.3073699"],"URL":"https:\/\/doi.org\/10.1145\/3072959.3073699","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,7,20]]},"assertion":[{"value":"2017-07-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}