{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T17:29:21Z","timestamp":1784568561068,"version":"3.55.0"},"reference-count":42,"publisher":"Springer Science and Business Media LLC","issue":"10-11","license":[{"start":{"date-parts":[[2020,6,11]],"date-time":"2020-06-11T00:00:00Z","timestamp":1591833600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,6,11]],"date-time":"2020-06-11T00:00:00Z","timestamp":1591833600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100000761","name":"Imperial College London","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100000761","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2020,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Image-to-image (i2i) translation is the dense regression problem of learning how to transform an input image into an output using aligned image pairs. Remarkable progress has been made in i2i translation with the advent of deep convolutional neural networks and particular using the learning paradigm of generative adversarial networks (GANs). In the absence of paired images, i2i translation is tackled with one or multiple domain transformations (i.e., CycleGAN, StarGAN etc.). In this paper, we study the problem of image-to-image translation, under a set of continuous parameters that correspond to a model describing a physical process. In particular, we propose the SliderGAN which transforms an input face image into a new one according to the continuous values of a statistical blendshape model of facial motion. We show that it is possible to edit a facial image according to expression and speech blendshapes, using sliders that control the continuous values of the blendshape model. This provides much more flexibility in various tasks, including but not limited to face editing, expression transfer and face neutralisation, comparing to models based on discrete expressions or action units.<\/jats:p>","DOI":"10.1007\/s11263-020-01338-7","type":"journal-article","created":{"date-parts":[[2020,6,11]],"date-time":"2020-06-11T06:02:36Z","timestamp":1591855356000},"page":"2629-2650","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":23,"title":["SliderGAN: Synthesizing Expressive Face Images by Sliding 3D Blendshape Parameters"],"prefix":"10.1007","volume":"128","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4345-1744","authenticated-orcid":false,"given":"Evangelos","family":"Ververas","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stefanos","family":"Zafeiriou","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,6,11]]},"reference":[{"key":"1338_CR1","unstructured":"Alami\u00a0Mejjati, Y., Richardt, C., Tompkin, J., Cosker, D., & Kim, K.I. (2018). Unsupervised attention-guided image-to-image translation (pp. 3693\u20133703)."},{"key":"1338_CR2","unstructured":"Amos, B., Ludwiczuk, B., & Satyanarayanan, M. (2016). Openface: A general-purpose face recognition library with mobile applications. Technical report CMU-CS-16-118, CMU School of Computer Science."},{"key":"1338_CR3","unstructured":"Arjovsky, M., Chintala, S., Bottou, L. (2017) Wasserstein generative adversarial networks. In Proceedings of the 34th international conference on machine learning, ICML 2017, Sydney, NSW, Australia, 6\u201311 August 2017, pp. 214\u2013223."},{"issue":"1","key":"1338_CR4","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1561\/2200000015","volume":"4","author":"F Bach","year":"2012","unstructured":"Bach, F., Jenatton, R., Mairal, J., & Obozinski, G. (2012). Optimization with sparsity-inducing penalties. Foundations and Trends in Machine Learning, 4(1), 1\u2013106.","journal-title":"Foundations and Trends in Machine Learning"},{"key":"1338_CR5","doi-asserted-by":"crossref","unstructured":"Benitez-Quiroz, C. F., Srinivasan, R., & Martinez, A. M. (2016). Emotionet: An accurate, real-time algorithm for the automatic annotation of a million facial expressions in the wild. In 2016 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 5562\u20135570).","DOI":"10.1109\/CVPR.2016.600"},{"key":"1338_CR6","doi-asserted-by":"crossref","unstructured":"Benitez-Quiroz, C. F., Wang, Y., & Martinez, A. M. (2017) Recognition of action units in the wild with deep nets and a new global-local loss. In ICCV (pp. 3990\u20133999).","DOI":"10.1109\/ICCV.2017.428"},{"key":"1338_CR7","first-page":"1683","volume":"29","author":"F Benitez-Quiroz","year":"2018","unstructured":"Benitez-Quiroz, F., Srinivasan, R., & Martinez, A. M. (2018). Discriminant functional learning of color features for the recognition of facial action units and their intensities. IEEE Transactions on Pattern Analysis and Machine Intelligence, 29, 1683.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1338_CR8","doi-asserted-by":"crossref","unstructured":"Booth, J., Antonakos, E., Ploumpis, S., Trigeorgis, G., Panagakis, Y., & Zafeiriou, S., et\u00a0al. (2017). 3d face morphable models \u201cin-the-wild\u201d. In Proceedings of the IEEE conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2017.580"},{"key":"1338_CR9","doi-asserted-by":"publisher","first-page":"2638","DOI":"10.1109\/TPAMI.2018.2832138","volume":"40","author":"J Booth","year":"2018","unstructured":"Booth, J., Roussos, A., Ververas, E., Antonakos, E., Poumpis, S., Panagakis, Y., et al. (2018). 3d reconstruction of \u201cin-the-wild\u201d faces in images and videos. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40, 2638.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1338_CR10","doi-asserted-by":"crossref","unstructured":"Booth, J., Roussos, A., Zafeiriou, S., Ponniahy, A., & Dunaway, D.: A 3d morphable model learnt from 10,000 faces. In 2016 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 5543\u20135552).","DOI":"10.1109\/CVPR.2016.598"},{"key":"1338_CR11","doi-asserted-by":"crossref","unstructured":"Cheng, S., Kotsia, I., Pantic, M., & Zafeiriou, S.: 4dfab: A large scale 4d database for facial expression analysis and biometric applications. In 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2018). Salt Lake City, Utah, US.","DOI":"10.1109\/CVPR.2018.00537"},{"key":"1338_CR12","unstructured":"Choi, Y., Choi, M., Kim, M., Ha, J. W., Kim, S., & Choo, J.: Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In The IEEE conference on computer vision and pattern recognition (CVPR)."},{"key":"1338_CR13","unstructured":"Chung, J. S., & Zisserman, A.: Lip reading in the wild. In Asian conference on computer vision."},{"key":"1338_CR14","doi-asserted-by":"crossref","unstructured":"Deng, J., Guo, J., & Zafeiriou, S. (2018). Arcface: Additive angular margin loss for deep face recognition. arXiv:1801.07698","DOI":"10.1109\/CVPR.2019.00482"},{"key":"1338_CR15","unstructured":"Ekman, P. (2002). Facial action coding system (facs). A human face."},{"issue":"38","key":"1338_CR16","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3306346.3323028","volume":"2019","author":"O Fried","year":"2019","unstructured":"Fried, O., Tewari, A., Zollh\u00f6fer, M., Finkelstein, A., Shechtman, E., Goldman, D. B., et al. (2019). Text-based editing of talking-head video. ACM Transactions on Graphics, 2019(38), 1\u201314.","journal-title":"ACM Transactions on Graphics"},{"key":"1338_CR17","doi-asserted-by":"crossref","unstructured":"Geng, Z., Cao, C., & Tulyakov, S. (2019) 3d guided fine-grained face manipulation.","DOI":"10.1109\/CVPR.2019.01005"},{"key":"1338_CR18","first-page":"2672","volume-title":"Advances in neural information processing system","author":"I Goodfellow","year":"2014","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., et al. (2014). Generative adversarial nets. In Z. Ghahramani, M. Welling, C. Cortes, N. D. Lawrence, & K. Q. Weinberger (Eds.), Advances in neural information processing system (pp. 2672\u20132680). Red Hook: Curran Associates Inc."},{"key":"1338_CR19","first-page":"5767","volume-title":"Advances in neural information processing systems","author":"I Gulrajani","year":"2017","unstructured":"Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., & Courville, A. C. (2017). Improved training of wasserstein gans. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, & R. Garnett (Eds.), Advances in neural information processing systems (Vol. 30, pp. 5767\u20135777). Red Hook: Curran Associates Inc."},{"key":"1338_CR20","unstructured":"Isola, P., Zhu, J. Y., Zhou, T., Efros, A. A. (204). Image-to-image translation with conditional adversarial networks. In CVPR."},{"key":"1338_CR21","unstructured":"Jolicoeur-Martineau, A. (2019). The relativistic discriminator: A key element missing from standard GAN. In International conference on learning representations."},{"key":"1338_CR22","doi-asserted-by":"crossref","unstructured":"Kim, H., Garrido, P., Tewari, A., Xu, W., Thies, J., Nie\u00dfner, N., et al. (2018). Deep video portraits. In ACM transactions on graphics 2018 (TOG).","DOI":"10.1145\/3197517.3201283"},{"key":"1338_CR23","unstructured":"Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. CoRR arXiv:1412.6980."},{"key":"1338_CR24","unstructured":"Li, M., Zuo, W., & Zhang, D. (2016). Deep identity-aware transfer of facial attributes. CoRR arXiv:1610.05586."},{"key":"1338_CR25","doi-asserted-by":"crossref","unstructured":"Li, S., Deng, W., & Du, J.: Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild. In 2017 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 2584\u20132593). IEEE.","DOI":"10.1109\/CVPR.2017.277"},{"issue":"6","key":"1338_CR26","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1145\/2508363.2508417","volume":"32","author":"T Neumann","year":"2013","unstructured":"Neumann, T., Varanasi, K., Wenger, S., Wacker, M., Magnor, M., & Theobalt, C. (2013). Sparse localized deformation components. ACM Transactions on Graphics (TOG), 32(6), 179.","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"1338_CR27","unstructured":"Perarnau, G., van\u00a0de Weijer, J., Raducanu, B., & \u00c1lvarez, J. M.: Invertible conditional gans for image editing. In it CoRR arXiv:1611.06355."},{"key":"1338_CR28","doi-asserted-by":"crossref","unstructured":"Pumarola, A., Agudo, A., Martinez, A., Sanfeliu, A., & Moreno-Noguer, F. (2018). Ganimation: Anatomically-aware facial animation from a single image. In Proceedings of the European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-01249-6_50"},{"key":"1338_CR29","doi-asserted-by":"crossref","unstructured":"Richardson, E., Sela, M., Or-El, R., & Kimmel, R. (2017). Learning detailed face reconstruction from a single image. In 2017 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 5553\u20135562). IEEE.","DOI":"10.1109\/CVPR.2017.589"},{"issue":"4","key":"1338_CR30","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1145\/3072959.3073640","volume":"36","author":"S Suwajanakorn","year":"2017","unstructured":"Suwajanakorn, S., Seitz, S. M., & Kemelmacher-Shlizerman, I. (2017). Synthesizing obama: Learning lip sync from audio. ACM Transactions on Graphics (TOG), 36(4), 95.","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"1338_CR31","doi-asserted-by":"crossref","unstructured":"Tewari, A., Zollh\u00f6fer, M., Garrido, P., Bernard, F., Kim, H., P\u00e9rez, P., & Theobalt, C. (2017). Self-supervised multi-level face model learning for monocular reconstruction at over 250 hz, vol. 2. arXiv preprint arXiv:1712.02859.","DOI":"10.1109\/CVPR.2018.00270"},{"key":"1338_CR32","doi-asserted-by":"crossref","unstructured":"Thies, J., Zollh\u00f6fer, M., Stamminger, M., Theobalt, C., & Nie\u00dfner, M. (2016). Face2Face: Real-time face capture and reenactment of RGB videos. In Proceedings computer vision and pattern recognition (CVPR) IEEE.","DOI":"10.1109\/CVPR.2016.262"},{"key":"1338_CR33","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3306346.3323035","volume":"38","author":"J Thies","year":"2019","unstructured":"Thies, J., Zollhhfer, M., & Niener, M. (2019). Deferred neural rendering: Image synthesis using neural textures. ACM Transactions on Graphics, 38, 1\u201312.","journal-title":"ACM Transactions on Graphics"},{"key":"1338_CR34","doi-asserted-by":"crossref","unstructured":"Tran, L., & Liu, X. (2018). Nonlinear 3d face morphable model. arXiv preprint arXiv:1804.03786","DOI":"10.1109\/CVPR.2018.00767"},{"key":"1338_CR35","unstructured":"Tzirakis, P., Papaioannou, A., Lattas, A., Tarasiou, M., Schuller, B., & Zafeiriou, S.: Synthesising 3d facial motion from in-the-wild speech. arXiv preprint arXiv:1904.07002."},{"key":"1338_CR36","unstructured":"Usman, B., Dufour, N., Saenko, K., & Bregler, C.: Puppetgan: Cross-domain image manipulation by demonstration. In The IEEE international conference on computer vision (ICCV)."},{"issue":"8","key":"1338_CR37","doi-asserted-by":"publisher","first-page":"1334","DOI":"10.1109\/TPAMI.2005.165","volume":"27","author":"L Wang","year":"2005","unstructured":"Wang, L., Zhang, Y., & Feng, J. (2005). On the euclidean distance of images. The IEEE Transactions on Pattern Analysis and Machine Intelligence, 27(8), 1334\u20131339.","journal-title":"The IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1338_CR38","doi-asserted-by":"crossref","unstructured":"Wang, X., Yu, K., Wu, S., Gu, J., Liu, Y., & Dong, C., et al. (2018). Enhanced super-resolution generative adversarial networks. In The European conference on computer vision workshops (ECCVW).","DOI":"10.1007\/978-3-030-11021-5_5"},{"key":"1338_CR39","doi-asserted-by":"crossref","unstructured":"Wiles, O., Koepke, A., & Zisserman, A. (2018). X2face: A network for controlling face generation by using images, audio, and pose codes. In European conference on computer vision.","DOI":"10.1007\/978-3-030-01261-8_41"},{"key":"1338_CR40","doi-asserted-by":"crossref","unstructured":"Wiles, O., Koepke, A. S., & Zisserman, A. (2018). X2face: A network for controlling face generation using images, audio, and pose codes. In Proceedings of the ECCV.","DOI":"10.1007\/978-3-030-01261-8_41"},{"issue":"7","key":"1338_CR41","doi-asserted-by":"publisher","first-page":"2479","DOI":"10.1109\/TSP.2009.2016892","volume":"57","author":"SJ Wright","year":"2009","unstructured":"Wright, S. J., Nowak, R. D., & Figueiredo, M. A. T. (2009). Sparse reconstruction by separable approximation. IEEE Transactions on Signal Processing, 57(7), 2479\u20132493.","journal-title":"IEEE Transactions on Signal Processing"},{"key":"1338_CR42","doi-asserted-by":"crossref","unstructured":"Zhu, J. Y., Park, T., Isola, P., & Efros, A. A. (2017). Unpaired image-to-image translation using cycle-consistent adversarial networkss. In 2017 IEEE international conference on computer vision (ICCV).","DOI":"10.1109\/ICCV.2017.244"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-020-01338-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-020-01338-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-020-01338-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,6,10]],"date-time":"2021-06-10T23:29:11Z","timestamp":1623367751000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-020-01338-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,6,11]]},"references-count":42,"journal-issue":{"issue":"10-11","published-print":{"date-parts":[[2020,11]]}},"alternative-id":["1338"],"URL":"https:\/\/doi.org\/10.1007\/s11263-020-01338-7","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,6,11]]},"assertion":[{"value":"15 May 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 May 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 June 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}