{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,21]],"date-time":"2026-05-21T03:44:54Z","timestamp":1779335094370,"version":"3.51.4"},"reference-count":69,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,8,7]],"date-time":"2024-08-07T00:00:00Z","timestamp":1722988800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,8,7]],"date-time":"2024-08-07T00:00:00Z","timestamp":1722988800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Intell."],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Conversation is an essential component of virtual avatar activities in the metaverse. With the development of natural language processing, significant breakthroughs have been made in text and voice conversation generation. However, face-to-face conversations account for the vast majority of daily conversations, while most existing methods focused on single-person talking head generation. In this work, we take a step further and consider generating realistic face-to-face conversation videos. Conversation generation is more challenging than single-person talking head generation, because it requires not only the generation of photo-realistic individual talking heads, but also the listener\u2019s response to the speaker. In this paper, we propose a novel unified framework based on the neural radiance field (NeRF) to address these challenges. Specifically, we model both the speaker and the listener with a NeRF framework under different conditions to control individual expressions. The speaker is driven by the audio signal, while the response of the listener depends on both visual and acoustic information. In this way, face-to-face conversation videos are generated between human avatars, with all the interlocutors modeled within the same network. Moreover, to facilitate future research on this task, we also collected a new human conversation dataset containing 34 video clips. Quantitative and qualitative experiments evaluate our method in different aspects, e.g., image quality, pose sequence trend, and natural rendering of the scene in the generated videos. Experimental results demonstrate that the avatars in the resulting videos are able to carry on a realistic conversation, and maintain individual styles.<\/jats:p>","DOI":"10.1007\/s44267-024-00057-8","type":"journal-article","created":{"date-parts":[[2024,8,7]],"date-time":"2024-08-07T08:03:38Z","timestamp":1723017818000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":25,"title":["DialogueNeRF: towards realistic avatar face-to-face conversation video generation"],"prefix":"10.1007","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3209-8965","authenticated-orcid":false,"given":"Yichao","family":"Yan","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zanwei","family":"Zhou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zi","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingnan","family":"Gao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaokang","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,8,7]]},"reference":[{"key":"57_CR1","first-page":"1378","volume-title":"Proceedings of the 33rd international conference on machine learning","author":"A. Kumar","year":"2016","unstructured":"Kumar, A., Irsoy, O., Su, J., Bradbury, J., English, R., Pierce, B., et al. (2016). Ask me anything: dynamic memory networks for natural language processing. In Proceedings of the 33rd international conference on machine learning (pp. 1378\u20131387). Stroudsburg: International Machine Learning Society."},{"key":"57_CR2","doi-asserted-by":"publisher","first-page":"702","DOI":"10.18653\/v1\/D18-1075","volume-title":"Proceedings of the 2018 conference on empirical methods in natural language processing","author":"L. Luo","year":"2018","unstructured":"Luo, L., Xu, J., Lin, J., Zeng, Q., & Sun, X. (2018). An auto-encoder matching model for learning utterance-level semantic dependency in dialogue generation. In Proceedings of the 2018 conference on empirical methods in natural language processing (pp. 702\u2013707). Stroudsburg: ACL."},{"key":"57_CR3","first-page":"2193","volume-title":"Proceedings of the annual meeting of the association for computational linguistics","author":"Y. Wang","year":"2018","unstructured":"Wang, Y., Liu, C., Huang, M., & Nie, L. (2018). Learning to ask questions in open-domain conversational systems with typed decoders. In Proceedings of the annual meeting of the association for computational linguistics (pp. 2193\u20132203). New York: ACM."},{"key":"57_CR4","first-page":"125","volume-title":"Proceedings of the 9th ISCA speech synthesis workshop","author":"A. van den Oord","year":"2016","unstructured":"van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., et al. (2016). WaveNet: a generative model for raw audio. In Proceedings of the 9th ISCA speech synthesis workshop (p. 125). Sunnyvale: ISCA."},{"key":"57_CR5","first-page":"4006","volume-title":"Proceedings of the 18th annual conference of the international speech communication association","author":"Y. Wang","year":"2017","unstructured":"Wang, Y., Skerry-Ryan, R., Stanton, D., Wu, Y., Weiss, R. J., Jaitly, N., et al. (2017). Tacotron: towards end-to-end speech synthesis. In Proceedings of the 18th annual conference of the international speech communication association (pp. 4006\u20134010). Red Hook: Curran Associates."},{"key":"57_CR6","first-page":"923","volume-title":"Proceedings of the 10th international conference on language resources and evaluation","author":"P. Lison","year":"2016","unstructured":"Lison, P., & Tiedemann, J. (2016). OpenSubtitles2016: extracting large parallel corpora from movie and TV subtitles. In Proceedings of the 10th international conference on language resources and evaluation (pp. 923\u2013929). Paris: European Language Resources Association."},{"key":"57_CR7","doi-asserted-by":"publisher","first-page":"285","DOI":"10.18653\/v1\/W15-4640","volume-title":"Proceedings of the 16th annual meeting of the special interest group on discourse and dialogue","author":"R. Lowe","year":"2015","unstructured":"Lowe, R., Pow, N., Serban, I., & Pineau, J. (2015). The Ubuntu dialogue corpus: a large dataset for research in unstructured multi-turn dialogue systems. In Proceedings of the 16th annual meeting of the special interest group on discourse and dialogue (pp. 285\u2013294). Stroudsburg: ACL."},{"key":"57_CR8","first-page":"76","volume-title":"Proceedings of the 2nd workshop on cognitive modeling and computational linguistics","author":"C. Danescu-Niculescu-Mizil","year":"2011","unstructured":"Danescu-Niculescu-Mizil, C., & Lee, L. (2011). Chameleons in imagined conversations: a new approach to understanding coordination of linguistic style in dialogs. In F. Keller & D. Reitter (Eds.), Proceedings of the 2nd workshop on cognitive modeling and computational linguistics (pp. 76\u201387). Stroudsburg: ACL."},{"key":"57_CR9","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511620539","volume-title":"Using language","author":"H. H. Clark","year":"1996","unstructured":"Clark, H. H. (1996). Using language. Cambridge: Cambridge University Press."},{"issue":"12","key":"57_CR10","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0208030","volume":"13","author":"P. H\u00f6mke","year":"2018","unstructured":"H\u00f6mke, P., Holler, J., & Levinson, S. C. (2018). Eye blinks are perceived as communicative signals in human face-to-face interaction. PLoS ONE, 13(12), e0208030.","journal-title":"PLoS ONE"},{"key":"57_CR11","first-page":"10093","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"D. Cudeiro","year":"2019","unstructured":"Cudeiro, D., Bolkart, T., Laidlaw, C., Ranjan, A., & Black, M. J. (2019). Capture, learning, and synthesis of 3D speaking styles. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 10093\u201310103). Piscataway: IEEE."},{"issue":"6","key":"57_CR12","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3072959.3092817","volume":"36","author":"L. Hu","year":"2017","unstructured":"Hu, L., Saito, S., Wei, L., Nagano, K., Seo, J., Fursund, J., et al. (2017). Avatar digitization from a single image for real-time rendering. ACM Transactions on Graphics, 36(6), 1\u201314.","journal-title":"ACM Transactions on Graphics"},{"key":"57_CR13","first-page":"2328","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition workshops","author":"H. X. Pham","year":"2017","unstructured":"Pham, H. X., Cheung, S., & Pavlovic, V. (2017). Speech-driven 3D facial animation with implicit emotional awareness: a deep learning approach. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops (pp. 2328\u20132336). Piscataway: IEEE."},{"issue":"6","key":"57_CR14","first-page":"1","volume":"39","author":"Y. Zhou","year":"2020","unstructured":"Zhou, Y., Han, X., Shechtman, E., Echevarria, J., Kalogerakis, E., & Li, D. (2020). Makeittalk: speaker-aware talking-head animation. ACM Transactions on Graphics, 39(6), 1\u201315.","journal-title":"ACM Transactions on Graphics"},{"key":"57_CR15","first-page":"7824","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"L. Chen","year":"2019","unstructured":"Chen, L., Maddox, R. K., Duan, Z., & Xu, C. (2019). Hierarchical cross-modal talking face generation with dynamic pixel-wise loss. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 7824\u20137833). Piscataway: IEEE."},{"key":"57_CR16","first-page":"716","volume-title":"Proceedings of the 16th European conference on computer vision","author":"J. Thies","year":"2020","unstructured":"Thies, J., Elgharib, M., Tewari, A., Theobalt, C., & Nie\u00dfner, M. (2020). Neural voice puppetry: audio-driven facial reenactment. In A. Vedaldi, H. Bischof, T. Brox, et al.(Eds.), Proceedings of the 16th European conference on computer vision (pp. 716\u2013731). Cham: Springer."},{"key":"57_CR17","doi-asserted-by":"crossref","unstructured":"Guo, Y., Chen, K., Liang, S., Liu, Y., Bao, H., & Zhang, J. (2021). Ad-NeRF: audio driven neural radiance fields for talking head synthesis. arXiv preprint. arXiv:2103.11078.","DOI":"10.1109\/ICCV48922.2021.00573"},{"key":"57_CR18","first-page":"8649","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"G. Gafni","year":"2021","unstructured":"Gafni, G., Thies, J., Zollh\u00f6fer, M., & Nie\u00dfner, M. (2021). Dynamic neural radiance fields for monocular 4D facial avatar reconstruction. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 8649\u20138658). Piscataway: IEEE."},{"issue":"1s","key":"57_CR19","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3176649","volume":"14","author":"G. G. Demisse","year":"2018","unstructured":"Demisse, G. G., Aouada, D., & Ottersten, B. E. (2018). Deformation-based 3D facial expression representation. ACM Transactions on Multimedia Computing Communications and Applications, 14(1s), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing Communications and Applications"},{"issue":"3","key":"57_CR20","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3571857","volume":"19","author":"H. Xue","year":"2023","unstructured":"Xue, H., Ling, J., Tang, A., Song, L., Xie, R., & Zhang, W. (2023). High-fidelity face reenactment via identity-matched correspondence learning. ACM Transactions on Multimedia Computing Communications and Applications, 19(3), 1\u201323.","journal-title":"ACM Transactions on Multimedia Computing Communications and Applications"},{"key":"57_CR21","doi-asserted-by":"publisher","first-page":"1145","DOI":"10.1109\/TIP.2023.3240835","volume":"32","author":"W. Yang","year":"2023","unstructured":"Yang, W., Chen, Z., Chen, C., Chen, G., & Wong, K. K. (2023). Deep face video inpainting via UV mapping. IEEE Transactions on Image Processing, 32, 1145\u20131157.","journal-title":"IEEE Transactions on Image Processing"},{"key":"57_CR22","first-page":"1153","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"A. Richard","year":"2021","unstructured":"Richard, A., Zollh\u00f6fer, M., Wen, Y., la Torre, F. D., & Sheikh, Y. (2021). Meshtalk: 3D face animation from speech using cross-modality disentanglement. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 1153\u20131162). Piscataway: IEEE."},{"issue":"1s","key":"57_CR23","first-page":"1","volume":"14","author":"S. Berretti","year":"2018","unstructured":"Berretti, S., Daoudi, M., Turaga, P. K., & Basu, A. (2018). Representation, analysis, and recognition of 3D humans: a survey. ACM Transactions on Multimedia Computing Communications and Applications, 14(1s), 1\u201336.","journal-title":"ACM Transactions on Multimedia Computing Communications and Applications"},{"issue":"1","key":"57_CR24","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3287309","volume":"15","author":"P. Pala","year":"2019","unstructured":"Pala, P., & Berretti, S. (2019). Reconstructing 3D face models by incremental aggregation and refinement of depth frames. ACM Transactions on Multimedia Computing Communications and Applications, 15(1), 1\u201324.","journal-title":"ACM Transactions on Multimedia Computing Communications and Applications"},{"issue":"4","key":"57_CR25","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3072959.3073640","volume":"36","author":"S. Suwajanakorn","year":"2017","unstructured":"Suwajanakorn, S., Seitz, S. M., & Kemelmacher-Shlizerman, I. (2017). Synthesizing obama: learning lip sync from audio. ACM Transactions on Graphics, 36(4), 1\u201313.","journal-title":"ACM Transactions on Graphics"},{"key":"57_CR26","first-page":"3661","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Z. Zhang","year":"2021","unstructured":"Zhang, Z., Li, L., Ding, Y., & Fan, C. (2021). Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 3661\u20133670). Piscataway: IEEE."},{"issue":"19","key":"57_CR27","doi-asserted-by":"publisher","DOI":"10.1007\/s44267-023-00021-y","volume":"1","author":"Z. Xu","year":"2023","unstructured":"Xu, Z., Shang, H., Yang, S., Xu, R., Yan, Y., Li, Y., et al. (2023). Hierarchical painter: Chinese landscape painting restoration with fine-grained styles. Visual Intelligence, 1(19), 19.","journal-title":"Visual Intelligence"},{"issue":"4","key":"57_CR28","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3528223.3530164","volume":"41","author":"R. Gal","year":"2022","unstructured":"Gal, R., Patashnik, O., Maron, H., Bermano, A. H., Chechik, G., & Cohen-Or, D. (2022). StyleGAN-NADA: clip-guided domain adaptation of image generators. ACM Transactions on Graphics, 41(4), 1\u201313.","journal-title":"ACM Transactions on Graphics"},{"key":"57_CR29","first-page":"852","volume-title":"Proceedings of the 35th international conference on neural information processing systems","author":"T. Karras","year":"2021","unstructured":"Karras, T., Aittala, M., Laine, S., H\u00e4rk\u00f6nen, E., Hellsten, J., Lehtinen, J., et al. (2021). Alias-free generative adversarial networks. In Proceedings of the 35th international conference on neural information processing systems (pp. 852\u2013863). Red Hook: Curran Associates."},{"key":"57_CR30","first-page":"4401","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"T. Karras","year":"2019","unstructured":"Karras, T., Laine, S., & Aila, T. (2019). A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 4401\u20134410). Piscataway: IEEE."},{"key":"57_CR31","first-page":"405","volume-title":"Proceedings of the 16th European conference on computer vision","author":"B. Mildenhall","year":"2020","unstructured":"Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., & Ng, R. (2020). NeRF: representing scenes as neural radiance fields for view synthesis. In A. Vedaldi, H. Bischof, T. Brox, et al.(Eds.), Proceedings of the 16th European conference on computer vision (pp. 405\u2013421). Cham: Springer."},{"key":"57_CR32","unstructured":"Yan, Y., Zhou, Z., Wang, Z., Gao, J., & Yang, X. (2023). DialogueNeRF: towards realistic avatar face-to-face conversation video generation. arXiv preprint. arXiv:2203.07931."},{"issue":"2","key":"57_CR33","doi-asserted-by":"publisher","first-page":"356","DOI":"10.1109\/TASL.2011.2125954","volume":"20","author":"X. Anguera","year":"2012","unstructured":"Anguera, X., Bozonnet, S., Evans, N., Fredouille, C., Friedland, G., & Vinyals, O. (2012). Speaker diarization: a review of recent research. IEEE Transactions on Audio, Speech, and Language Processing, 20(2), 356\u2013370.","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"key":"57_CR34","first-page":"1368","volume-title":"Proceedings of the 19th annual conference of the international speech communication association","author":"P.-A. Broux","year":"2018","unstructured":"Broux, P.-A., Desnous, F., Larcher, A., Petitrenaud, S., Carrive, J., & Meignier, S. (2018). S4D: speaker diarization toolkit in Python. In Proceedings of the 19th annual conference of the international speech communication association (pp. 1368\u20131372). Red Hook: Curran Associates."},{"key":"57_CR35","first-page":"6465","volume-title":"Proceedings of the IEEE international conference on acoustics, speech and signal processing","author":"D. Povey","year":"2011","unstructured":"Povey, D., Ghoshal, A., Boulianne, G., Burget, L., Glembek, O., Goel, N., et al. (2011). The Pytorch-kaldi speech recognition toolkit. In Proceedings of the IEEE international conference on acoustics, speech and signal processing (pp. 6465\u20136469). Piscataway: IEEE."},{"key":"57_CR36","first-page":"737","volume-title":"Proceedings of the IEEE international conference on acoustics, speech, and signal processing","author":"J.-F. Bonastre","year":"2005","unstructured":"Bonastre, J.-F., Wils, F., & Meignier, S. (2005). Alize, a free toolkit for speaker recognition. In Proceedings of the IEEE international conference on acoustics, speech, and signal processing (pp. 737\u2013740). Piscataway: IEEE."},{"issue":"12","key":"57_CR37","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0144610","volume":"10","author":"T. Giannakopoulos","year":"2015","unstructured":"Giannakopoulos, T. (2015). Pyaudioanalysis: an open-source python library for audio signal analysis. PLoS ONE, 10(12), e0144610.","journal-title":"PLoS ONE"},{"key":"57_CR38","first-page":"7124","volume-title":"Proceedings of the IEEE international conference on acoustics, speech and signal processing","author":"H. Bredin","year":"2020","unstructured":"Bredin, H., Yin, R., Coria, J. M., Gelly, G., Korshunov, P., Lavechin, M., et al. (2020). Pyannote. Audio: neural building blocks for speaker diarization. In Proceedings of the IEEE international conference on acoustics, speech and signal processing (pp. 7124\u20137128). Piscataway: IEEE."},{"key":"57_CR39","first-page":"1119","volume-title":"Proceedings of the 33rd international conference on neural information processing systems","author":"V. Sitzmann","year":"2019","unstructured":"Sitzmann, V., Zollh\u00f6fer, M., & Wetzstein, G. (2019). Scene representation networks: continuous 3D-structure-aware neural scene representations. In H. M. Wallach, H. Larochelle, A. Beygelzimer, et al.(Eds.), Proceedings of the 33rd international conference on neural information processing systems (pp. 1119\u20131130). Red Hook: Curran Associates."},{"key":"57_CR40","first-page":"20154","volume-title":"Proceedings of the 34th international conference on neural information processing systems","author":"K. Schwarz","year":"2020","unstructured":"Schwarz, K., Liao, Y., Niemeyer, M., & Geiger, A. (2020). GRAF: generative radiance fields for 3D-aware image synthesis. In Proceedings of the 34th international conference on neural information processing systems (pp. 20154\u201320166). Red Hook: Curran Associates."},{"issue":"3","key":"57_CR41","doi-asserted-by":"publisher","first-page":"1827","DOI":"10.1109\/TCSVT.2023.3298929","volume":"34","author":"R. Li","year":"2024","unstructured":"Li, R., Dai, P., Liu, G., Zhang, S., Zeng, B., & Liu, S. (2024). PBR-GAN: imitating physically based rendering with generative adversarial networks. IEEE Transactions on Circuits and Systems for Video Technology, 34(3), 1827\u20131840.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"57_CR42","first-page":"11453","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"M. Niemeyer","year":"2021","unstructured":"Niemeyer, M., & Geiger, A. (2021). Giraffe: representing scenes as compositional generative neural feature fields. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 11453\u201311464). Piscataway: IEEE."},{"key":"57_CR43","first-page":"9054","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"S. Peng","year":"2021","unstructured":"Peng, S., Zhang, Y., Xu, Y., Wang, Q., Shuai, Q., Bao, H., et al. (2021). Neural body: implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 9054\u20139063). Piscataway: IEEE."},{"issue":"6","key":"57_CR44","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3478513.3480487","volume":"40","author":"K. Park","year":"2021","unstructured":"Park, K., Sinha, U., Hedman, P., Barron, J. T., Bouaziz, S., Goldman, D. B., et al. (2021). HyperNeRF: a higher-dimensional representation for topologically varying neural radiance fields. ACM Transactions on Graphics, 40(6), 1\u201312.","journal-title":"ACM Transactions on Graphics"},{"key":"57_CR45","first-page":"520","volume-title":"Proceedings of the 15th European conference on computer vision","author":"L. Chen","year":"2018","unstructured":"Chen, L., Li, Z., Maddox, R. K., Duan, Z., & Xu, C. (2018). Lip movements generation at a glance. In V. Ferrari, M. Hebert, & C. Sminchisescu (Eds.), Proceedings of the 15th European conference on computer vision (pp. 520\u2013535). Cham: Springer."},{"key":"57_CR46","first-page":"787","volume-title":"Proceedings of the IEEE international conference on data mining","author":"L. Yu","year":"2019","unstructured":"Yu, L., Yu, J., & Ling, Q. (2019). Mining audio, text and visual information for talking face generation. In Proceedings of the IEEE international conference on data mining (pp. 787\u2013795). Piscataway: IEEE."},{"key":"57_CR47","first-page":"484","volume-title":"Proceedings of the ACM international conference on multimedia","author":"K. R. Prajwal","year":"2020","unstructured":"Prajwal, K. R., Mukhopadhyay, R., Namboodiri, V. P., & Jawahar, C. (2020). A lip sync expert is all you need for speech to lip generation in the wild. In Proceedings of the ACM international conference on multimedia (pp. 484\u2013492). New York: ACM."},{"key":"57_CR48","first-page":"919","volume-title":"Proceedings of the 28th international joint conference on artificial intelligence","author":"Y. Song","year":"2019","unstructured":"Song, Y., Zhu, J., Li, D., Wang, X., & Qi, H. (2019). Talking face generation by conditional recurrent adversarial network. In Proceedings of the 28th international joint conference on artificial intelligence (pp. 919\u2013925). Cham: Springer."},{"issue":"5","key":"57_CR49","doi-asserted-by":"publisher","first-page":"1398","DOI":"10.1007\/s11263-019-01251-8","volume":"128","author":"K. Vougioukas","year":"2020","unstructured":"Vougioukas, K., Petridis, S., & Pantic, M. (2020). Realistic speech-driven facial animation with GANs. International Journal of Computer Vision, 128(5), 1398\u20131413.","journal-title":"International Journal of Computer Vision"},{"issue":"3","key":"57_CR50","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3571746","volume":"19","author":"S. Liu","year":"2023","unstructured":"Liu, S., & Wang, H. (2023). Talking face generation via facial anatomy. ACM Transactions on Multimedia Computing Communications and Applications, 19(3), 1\u201319.","journal-title":"ACM Transactions on Multimedia Computing Communications and Applications"},{"issue":"4","key":"57_CR51","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3072959.3073658","volume":"36","author":"T. Karras","year":"2017","unstructured":"Karras, T., Aila, T., Laine, S., Herva, A., & Lehtinen, J. (2017). Audio-driven facial animation by joint end-to-end learning of pose and emotion. ACM Transactions on Graphics, 36(4), 1\u201312.","journal-title":"ACM Transactions on Graphics"},{"key":"57_CR52","first-page":"173","volume-title":"Proceedings of the international conference on machine learning","author":"D. Amodei","year":"2016","unstructured":"Amodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., et al. (2016). Deep speech 2: end-to-end speech recognition in English and mandarin. In Proceedings of the international conference on machine learning (pp. 173\u2013182). Stroudsburg: International Machine Learning Society."},{"key":"57_CR53","unstructured":"Bai, S., Kolter, J. Z., & Koltun, V. (2018). An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint. arXiv:1803.01271."},{"key":"57_CR54","first-page":"10934","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"S. Ploumpis","year":"2019","unstructured":"Ploumpis, S., Wang, H., Pears, N., Smith, W. A., & Zafeiriou, S. (2019). Combining 3D morphable models: a large scale face-and-head model. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 10934\u201310943). Piscataway: IEEE."},{"key":"57_CR55","unstructured":"Yao, S., Zhong, R., Yan, Y., Zhai, G., & Yang, X. (2022). DFA-NeRF: personalized talking head generation via disentangled face attributes neural rendering. arXiv preprint. arXiv:2201.00791."},{"key":"57_CR56","first-page":"5865","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"K. Park","year":"2021","unstructured":"Park, K., Sinha, U., Barron, J. T., Bouaziz, S., Goldman, D. B., Seitz, S. M., et al. (2021). Nerfies: deformable neural radiance fields. In Proceedings of the IEEE\/CVF international conference on computer vision (pp. 5865\u20135874). Piscataway: IEEE."},{"key":"57_CR57","first-page":"87","volume-title":"Proceedings of the international workshop on quality of multimedia experience","author":"N. Narvekar","year":"2009","unstructured":"Narvekar, N., & Karam, L. (2009). A no-reference perceptual image sharpness metric based on a cumulative probability of blur detection. In Proceedings of the international workshop on quality of multimedia experience (pp. 87\u201391). Piscataway: IEEE."},{"key":"57_CR58","first-page":"2366","volume-title":"Proceedings of the international conference on pattern recognition","author":"A. Hore","year":"2010","unstructured":"Hore, A., & Ziou, D. (2010). Image quality metrics: PSNR vs. SSIM. In Proceedings of the international conference on pattern recognition (pp. 2366\u20132369). Piscataway: IEEE."},{"issue":"4","key":"57_CR59","doi-asserted-by":"publisher","first-page":"600","DOI":"10.1109\/TIP.2003.819861","volume":"13","author":"Z. Wang","year":"2004","unstructured":"Wang, Z., Bovik, A. C., Sheikh, H. R., & Simoncelli, E. P. (2004). Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 13(4), 600\u2013612.","journal-title":"IEEE Transactions on Image Processing"},{"key":"57_CR60","first-page":"251","volume-title":"Proceedings of the Asian conference on computer vision workshops","author":"J. S. Chung","year":"2016","unstructured":"Chung, J. S., & Zisserman, A. (2016). Out of time: automated lip sync in the wild. In Proceedings of the Asian conference on computer vision workshops (pp. 251\u2013263). Cham: Springer."},{"key":"57_CR61","first-page":"8026","volume-title":"Proceedings of the 33rd international conference on neural information processing systems","author":"A. Paszke","year":"2019","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., et al. (2019). Pytorch: an imperative style, high-performance deep learning library. In Proceedings of the 33rd international conference on neural information processing systems (pp. 8026\u20138037). Red Hook: Curran Associates."},{"key":"57_CR62","volume-title":"Proceedings of the 3rd international conference on learning representations, San Diego, USA","author":"D. P. Kingma","year":"2015","unstructured":"Kingma, D. P., & Ba, J. (2015). Adam: a method for stochastic optimization. [Poster presentation]. Proceedings of the 3rd international conference on learning representations, San Diego, USA."},{"key":"57_CR63","doi-asserted-by":"crossref","unstructured":"Lu, Y., Chai, J., & Cao, X. (2021). Live speech portraits: real-time photorealistic talking-head animation. arXiv preprint. arXiv:2109.10595.","DOI":"10.1145\/3478513.3480484"},{"key":"57_CR64","first-page":"4176","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"H. Zhou","year":"2021","unstructured":"Zhou, H., Sun, Y., Wu, W., Loy, C. C., Wang, X., & Liu, Z. (2021). Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4176\u20134186). Piscataway: IEEE."},{"key":"57_CR65","first-page":"500","volume-title":"Proceedings of the IEEE international symposium on multimedia","author":"J. Xu","year":"2011","unstructured":"Xu, J., Xing, L., Perkis, A., & Jiang, Y. (2011). On the properties of mean opinion scores for quality of experience management. In Proceedings of the IEEE international symposium on multimedia (pp. 500\u2013505). Piscataway: IEEE."},{"key":"57_CR66","first-page":"14326","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"S. J. Garbin","year":"2021","unstructured":"Garbin, S. J., Kowalski, M., Johnson, M., Shotton, J., & Valentin, J. (2021). FastNeRF: high-fidelity neural rendering at 200FPS. In Proceedings of the IEEE international conference on computer vision (pp. 14326\u201314335). Piscataway: IEEE."},{"key":"57_CR67","first-page":"5732","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"A. Yu","year":"2021","unstructured":"Yu, A., Li, R., Tancik, M., Li, H., Ng, R., & Kanazawa, A. (2021). Plenoctrees for real-time rendering of neural radiance fields. In Proceedings of the IEEE international conference on computer vision (pp. 5732\u20135741). Piscataway: IEEE."},{"issue":"4","key":"57_CR68","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3450626.3459863","volume":"40","author":"S. Lombardi","year":"2021","unstructured":"Lombardi, S., Simon, T., Schwartz, G., Zollhoefer, M., Sheikh, Y., & Saragih, J. (2021). Mixture of volumetric primitives for efficient neural rendering. ACM Transactions on Graphics, 40(4), 1\u201313.","journal-title":"ACM Transactions on Graphics"},{"issue":"4","key":"57_CR69","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3528223.3530127","volume":"41","author":"T. M\u00fcller","year":"2022","unstructured":"M\u00fcller, T., Evans, A., Schied, C., & Keller, A. (2022). Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics, 41(4), 1\u201315.","journal-title":"ACM Transactions on Graphics"}],"container-title":["Visual Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-024-00057-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44267-024-00057-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44267-024-00057-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,7]],"date-time":"2024-08-07T09:16:57Z","timestamp":1723022217000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44267-024-00057-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,7]]},"references-count":69,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["57"],"URL":"https:\/\/doi.org\/10.1007\/s44267-024-00057-8","relation":{},"ISSN":["2731-9008"],"issn-type":[{"value":"2731-9008","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,7]]},"assertion":[{"value":"7 January 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 July 2024","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 July 2024","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 August 2024","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"24"}}