{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,7]],"date-time":"2025-08-07T20:35:12Z","timestamp":1754598912377,"version":"3.41.0"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,12,16]],"date-time":"2024-12-16T00:00:00Z","timestamp":1734307200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"H2020 SPRING","award":["871245"],"award-info":[{"award-number":["871245"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,1,31]]},"abstract":"<jats:p>In this work, we address the task of unconditional head motion generation to animate still human faces in a low-dimensional semantic space from a single reference pose. Different from traditional audio-conditioned talking head generation that seldom puts emphasis on realistic head motions, we devise a GAN-based architecture that learns to synthesize rich head motion sequences over long duration while maintaining low error-accumulation levels. In particular, the autoregressive generation of incremental outputs ensures smooth trajectories, while a multi-scale discriminator on input pairs drives generation toward better handling of high- and low-frequency signals and less mode collapse. We experimentally demonstrate the relevance of the proposed method and show its superiority compared to models that attained state-of-the-art performances on similar tasks.<\/jats:p>","DOI":"10.1145\/3635154","type":"journal-article","created":{"date-parts":[[2023,12,6]],"date-time":"2023-12-06T12:24:13Z","timestamp":1701865453000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Autoregressive GAN for Semantic Unconditional Head Motion Generation"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0795-0415","authenticated-orcid":false,"given":"Louis","family":"Airale","sequence":"first","affiliation":[{"name":"Univ. Grenoble Alpes, CNRS, Grenoble INP, LIG, Montbonnot, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5354-1084","authenticated-orcid":false,"given":"Xavier","family":"Alameda-Pineda","sequence":"additional","affiliation":[{"name":"Univ. Grenoble Alpes, Inria, CNRS, Grenoble INP, LJK, Montbonnot, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6927-8930","authenticated-orcid":false,"given":"St\u00e9phane","family":"Lathuili\u00e8re","sequence":"additional","affiliation":[{"name":"LTCI, T\u00e9l\u00e9com Paris, Institut polytechnique de Paris, Palaiseau, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8825-0973","authenticated-orcid":false,"given":"Dominique","family":"Vaufreydaz","sequence":"additional","affiliation":[{"name":"Univ. Grenoble Alpes, CNRS, Grenoble INP, LIG, Saint-Martin-d'Heres, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,12,16]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/TAFFC.2022.3171719"},{"key":"e_1_3_2_3_2","first-page":"11333","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Aliakbarian Sadegh","year":"2021","unstructured":"Sadegh Aliakbarian, Fatemeh Saleh, Lars Petersson, Stephen Gould, and Mathieu Salzmann. 2021. Contextually plausible and diverse 3D human motion prediction. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 11333\u201311342."},{"key":"e_1_3_2_4_2","first-page":"214","volume-title":"Proceedings of the 34th International Conference on Machine Learning (Proceedings of Machine Learning Research)","volume":"70","author":"Arjovsky Martin","year":"2017","unstructured":"Martin Arjovsky, Soumith Chintala, and L\u00e9on Bottou. 2017. Wasserstein generative adversarial networks. In Proceedings of the 34th International Conference on Machine Learning (Proceedings of Machine Learning Research), Doina Precup and Yee Whye Teh (Eds.), Vol. 70. PMLR, 214\u2013223."},{"key":"e_1_3_2_5_2","unstructured":"Xiaoyu Bie Wen Guo Simon Leglaive Lauren Girin Francesc Moreno-Noguer and Xavier Alameda-Pineda. 2022. HiT-DVAE: Human motion generation via hierarchical transformer dynamical VAE. Retrieved from https:\/\/arXiv:2204.01565"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_7_2","unstructured":"Tong Che Yanran Li Athul Paul Jacob Yoshua Bengio and Wenjie Li. 2016. Mode regularized generative adversarial networks. Retrieved from https:\/\/arXiv:1612.02136"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58545-7_3"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00802"},{"key":"e_1_3_2_10_2","article-title":"InfoGAN: Interpretable representation learning by information maximizing generative adversarial nets","volume":"29","author":"Chen Xi","year":"2016","unstructured":"Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel. 2016. InfoGAN: Interpretable representation learning by information maximizing generative adversarial nets. Adv. Neural Info. Process. Syst. 29 (2016).","journal-title":"Adv. Neural Info. Process. Syst."},{"key":"e_1_3_2_11_2","volume-title":"Proceedings of the 19th Annual Conference of the International Speech Communication Association (INTERSPEECH\u201918)","author":"Chung J. S.","year":"2018","unstructured":"J. S. Chung, A. Nagrani, and A. Zisserman. 2018. VoxCeleb2: Deep speaker recognition. In Proceedings of the 19th Annual Conference of the International Speech Communication Association (INTERSPEECH\u201918)."},{"key":"e_1_3_2_12_2","first-page":"408","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Das Dipanjan","year":"2020","unstructured":"Dipanjan Das, Sandika Biswas, Sanjana Sinha, and Brojeshwar Bhowmick. 2020. Speech-driven facial animation using cascaded GANs for learning of motion and texture. In Proceedings of the European Conference on Computer Vision. Springer, 408\u2013424."},{"key":"e_1_3_2_13_2","volume-title":"Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL)."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01821"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2016.12.001"},{"key":"e_1_3_2_16_2","first-page":"2672","volume-title":"Proceedings of the Advances in Neural Information Processing Systems (NeurIPS\u201914)","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS\u201914). 2672\u20132680."},{"key":"e_1_3_2_17_2","doi-asserted-by":"crossref","unstructured":"David Greenwood Iain Matthews and Stephen Laycock. 2018. Joint learning of facial expression and head pose from speech. In Proceedings of the Annual Conference of the International Speech Communication Association (INTERSPEECH\u201918).","DOI":"10.21437\/Interspeech.2018-2587"},{"key":"e_1_3_2_18_2","first-page":"2255","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"Gupta Agrim","year":"2018","unstructured":"Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi. 2018. Social GAN: Socially acceptable trajectories with generative adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201918). 2255\u20132264."},{"key":"e_1_3_2_19_2","first-page":"6626","article-title":"GANs trained by a two time-scale update rule converge to a local Nash equilibrium","volume":"30","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. Adv. Neural Info. Process. Syst. 30 (2017), 6626\u20136637.","journal-title":"Adv. Neural Info. Process. Syst."},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.632"},{"key":"e_1_3_2_21_2","unstructured":"Xinya Ji Hang Zhou Kaisiyuan Wang Qianyi Wu Wayne Wu Feng Xu and Xun Cao. 2022. EAMM: One-shot emotional talking face via audio-based emotion-aware motion model. Retrieved from https:\/\/arXiv:2205.15278"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073658"},{"key":"e_1_3_2_23_2","article-title":"Auto-encoding variational Bayes","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Max Welling. 2014. Auto-encoding variational Bayes. In Proceedings of the 2nd International Conference on Learning Representations (ICLR\u201914).","journal-title":"Proceedings of the 2nd International Conference on Learning Representations (ICLR\u201914)"},{"key":"e_1_3_2_24_2","article-title":"MelGAN: Generative adversarial networks for conditional waveform synthesis","volume":"32","author":"Kumar Kundan","year":"2019","unstructured":"Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Br\u00e9bisson, Yoshua Bengio, and Aaron C. Courville. 2019. MelGAN: Generative adversarial networks for conditional waveform synthesis. Adv. Neural Info. Process. Syst. 32 (2019).","journal-title":"Adv. Neural Info. Process. Syst."},{"key":"e_1_3_2_25_2","first-page":"8553","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Kundu Jogendra Nath","year":"2019","unstructured":"Jogendra Nath Kundu, Maharshi Gor, and R. Venkatesh Babu. 2019. BIHMP-GAN: Bidirectional 3D human motion prediction GAN. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 8553\u20138560."},{"key":"e_1_3_2_26_2","unstructured":"Jae Hyun Lim and Jong Chul Ye. 2017. Geometric GAN. Retrieved from https:\/\/arXiv:1705.02894"},{"key":"e_1_3_2_27_2","unstructured":"Xiao Lin and Mohamed R. Amer. 2018. Human motion modeling using DVGANs. Retrieved from https:\/\/arXiv:1804.10652"},{"key":"e_1_3_2_28_2","article-title":"PacGAN: The power of two samples in generative adversarial networks","volume":"31","author":"Lin Zinan","year":"2018","unstructured":"Zinan Lin, Ashish Khetan, Giulia Fanti, and Sewoong Oh. 2018. PacGAN: The power of two samples in generative adversarial networks. Adv. Neural Info. Process. Syst. 31 (2018).","journal-title":"Adv. Neural Info. Process. Syst."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.497"},{"key":"e_1_3_2_30_2","first-page":"13829","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Meshry Moustafa","year":"2021","unstructured":"Moustafa Meshry, Saksham Suri, Larry S. Davis, and Abhinav Shrivastava. 2021. Learned spatial representations for few-shot talking-head synthesis. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 13829\u201313838."},{"key":"e_1_3_2_31_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Morrison Max","year":"2022","unstructured":"Max Morrison, Rithesh Kumar, Kundan Kumar, Prem Seetharaman, Aaron Courville, and Yoshua Bengio. 2022. Chunked autoregressive GAN for conditional waveform synthesis. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=v3aeIsY_vVX"},{"issue":"2","key":"e_1_3_2_32_2","doi-asserted-by":"crossref","first-page":"848","DOI":"10.1109\/TPAMI.2020.3002500","article-title":"Dynamic facial expression generation on hilbert hypersphere with conditional Wasserstein generative adversarial nets","volume":"44","author":"Otberdout Naima","year":"2020","unstructured":"Naima Otberdout, Mohamed Daoudi, Anis Kacem, Lahoucine Ballihi, and Stefano Berretti. 2020. Dynamic facial expression generation on hilbert hypersphere with conditional Wasserstein generative adversarial nets. IEEE Trans. Pattern Anal. Mach. Intell. 44, 2 (2020), 848\u2013863.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01080"},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","unstructured":"Micha\u0142 Stypu\u0142kowski Konstantinos Vougioukas Sen He Maciej Zi\u0119ba Stavros Petridis and Maja Pantic. 2023. Diffused heads: Diffusion models beat GANs on talking-face generation. Retrieved from https:\/\/arXiv:2301.03396","DOI":"10.1109\/WACV57701.2024.00502"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073640"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.308"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073699"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00165"},{"key":"e_1_3_2_39_2","unstructured":"Thomas Unterthiner Sjoerd van Steenkiste Karol Kurach Raphael Marinier Marcin Michalski and Sylvain Gelly. 2018. Towards accurate generative models of video: A new metric and challenges. Retrieved from https:\/\/arXiv:1812.01717"},{"key":"e_1_3_2_40_2","article-title":"Neural discrete representation learning","volume":"30","author":"Oord Aaron Van Den","year":"2017","unstructured":"Aaron Van Den Oord, Oriol Vinyals et\u00a0al. 2017. Neural discrete representation learning. Adv. Neural Info. Process. Syst. 30 (2017).","journal-title":"Adv. Neural Info. Process. Syst."},{"key":"e_1_3_2_41_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Adv. Neural Info. Process. Syst. 30 (2017).","journal-title":"Adv. Neural Info. Process. Syst."},{"key":"e_1_3_2_42_2","first-page":"3560","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Villegas Ruben","year":"2017","unstructured":"Ruben Villegas, Jimei Yang, Yuliang Zou, Sungryull Sohn, Xunyu Lin, and Honglak Lee. 2017. Learning to generate long-term future via hierarchical prediction. In Proceedings of the International Conference on Machine Learning. PMLR, 3560\u20133569."},{"key":"e_1_3_2_43_2","first-page":"613","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems (NeurIPS\u201916)","author":"Vondrick Carl","year":"2016","unstructured":"Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. 2016. Generating videos with scene dynamics. In Proceedings of the Conference on Advances in Neural Information Processing Systems (NeurIPS\u201916). 613\u2013621."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-019-01251-8"},{"key":"e_1_3_2_45_2","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI\u201921)","author":"Wang Suzhe","year":"2021","unstructured":"Suzhe Wang, Lincheng Li, Yu Ding, Changjie Fan, and Xin Yu. 2021. Audio2Head: Audio-driven one-shot talking-head generation with natural head motion. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI\u201921)."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i3.20154"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00917"},{"key":"e_1_3_2_48_2","unstructured":"Ran Yi Zipeng Ye Juyong Zhang Hujun Bao and Yong-Jin Liu. 2020. Audio-driven talking face video generation with learning-based personalized head pose. Retrieved from https:\/\/arXiv:2002.10137"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2020.2973374"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58610-2_31"},{"key":"e_1_3_2_51_2","first-page":"9459","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201919)","author":"Zakharov Egor","year":"2019","unstructured":"Egor Zakharov, Aliaksandra Shysheya, Egor Burkov, and Victor Lempitsky. 2019. Few-shot adversarial learning of realistic neural talking head models. In Proceedings of the IEEE International Conference on Computer Vision (ICCV\u201919). 9459\u20139468."},{"key":"e_1_3_2_52_2","first-page":"1991","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhao Ruiqi","year":"2021","unstructured":"Ruiqi Zhao, Tianyi Wu, and Guodong Guo. 2021. Sparse to dense motion transfer for face image animation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 1991\u20132000."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33019299"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00416"},{"issue":"6","key":"e_1_3_2_55_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3414685.3417774","article-title":"Makelttalk: Speaker-aware talking-head animation","volume":"39","author":"Zhou Yang","year":"2020","unstructured":"Yang Zhou, Xintong Han, Eli Shechtman, Jose Echevarria, Evangelos Kalogerakis, and Dingzeyu Li. 2020. Makelttalk: Speaker-aware talking-head animation. ACM Trans. Graph. 39, 6 (2020), 1\u201315.","journal-title":"ACM Trans. Graph."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3635154","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3635154","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:49:06Z","timestamp":1750286946000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3635154"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,16]]},"references-count":54,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1,31]]}},"alternative-id":["10.1145\/3635154"],"URL":"https:\/\/doi.org\/10.1145\/3635154","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2024,12,16]]},"assertion":[{"value":"2023-04-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-14","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-16","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}