{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T19:02:19Z","timestamp":1783105339654,"version":"3.54.6"},"reference-count":93,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2024,7,19]],"date-time":"2024-07-19T00:00:00Z","timestamp":1721347200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2024,7,19]]},"abstract":"<jats:p>Faithful real-time facial animation is essential for avatar-mediated telepresence in Virtual Reality (VR). To emulate authentic communication, avatar animation needs to be efficient and accurate: able to capture both extreme and subtle expressions within a few milliseconds to sustain the rhythm of natural conversations. The oblique and incomplete views of the face, variability in the donning of headsets, and illumination variation due to the environment are some of the unique challenges in generalization to unseen faces. In this paper, we present a method that can animate a photorealistic avatar in realtime from head-mounted cameras (HMCs) on a consumer VR headset. We present a self-supervised learning approach, based on a cross-view reconstruction objective, that enables generalization to unseen users. We present a lightweight expression calibration mechanism that increases accuracy with minimal additional cost to run-time efficiency. We present an improved parameterization for precise ground-truth generation that provides robustness to environmental variation. The resulting system produces accurate facial animation for unseen users wearing VR headsets in realtime. We compare our approach to prior face-encoding methods demonstrating significant improvements in both quantitative metrics and qualitative results.<\/jats:p>","DOI":"10.1145\/3658234","type":"journal-article","created":{"date-parts":[[2024,7,19]],"date-time":"2024-07-19T14:47:57Z","timestamp":1721400477000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Universal Facial Encoding of Codec Avatars from VR Headsets"],"prefix":"10.1145","volume":"43","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-7969-9429","authenticated-orcid":false,"given":"Shaojie","family":"Bai","sequence":"first","affiliation":[{"name":"Meta Reality Labs, Pittsburgh, Pennsylvania, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-1690-6928","authenticated-orcid":false,"given":"Te-Li","family":"Wang","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Pittsburgh, Pennsylvania, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-7055-5757","authenticated-orcid":false,"given":"Chenghui","family":"Li","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Pittsburgh, Pennsylvania, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-1010-6476","authenticated-orcid":false,"given":"Akshay","family":"Venkatesh","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Pittsburgh, Pennsylvania, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0972-7455","authenticated-orcid":false,"given":"Tomas","family":"Simon","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Pittsburgh, Pennsylvania, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7992-2814","authenticated-orcid":false,"given":"Chen","family":"Cao","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Pittsburgh, Pennsylvania, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8781-5573","authenticated-orcid":false,"given":"Gabriel","family":"Schwartz","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Pittsburgh, Pennsylvania, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6218-5029","authenticated-orcid":false,"given":"Jason","family":"Saragih","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Pittsburgh, Pennsylvania, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6512-9348","authenticated-orcid":false,"given":"Yaser","family":"Sheikh","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Pittsburgh, Pennsylvania, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0092-3441","authenticated-orcid":false,"given":"Shih-En","family":"Wei","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, Pittsburgh, Pennsylvania, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,7,19]]},"reference":[{"key":"e_1_2_2_1_1","first-page":"1","article-title":"Digital Ira: Creating a Real-time Photoreal Digital Actor. In ACM SIGGRAPH 2013 Posters (SIGGRAPH '13). ACM, New York","volume":"1","author":"Alexander Oleg","year":"2013","unstructured":"Oleg Alexander, Graham Fyffe, Jay Busch, Xueming Yu, Ryosuke Ichikari, Andrew Jones, Paul Debevec, Jorge Jimenez, Etienne Danvoye, Bernardo Antionazzi, Mike Eheler, Zybnek Kysela, and Javier von der Pahlen. 2013. Digital Ira: Creating a Real-time Photoreal Digital Actor. In ACM SIGGRAPH 2013 Posters (SIGGRAPH '13). ACM, New York, NY, USA, 1:1--1:1.","journal-title":"NY, USA"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/MCG.2010.65"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.0956-7976.2005.01548.x"},{"key":"e_1_2_2_4_1","unstructured":"Apple. 2024. Set up your Persona (beta) on Apple Vision Pro. https:\/\/support.apple.com\/en-us\/118496."},{"key":"e_1_2_2_5_1","volume-title":"A critical analysis of self-supervision, or what we can learn from a single image. arXiv preprint arXiv:1904.13132","author":"Asano Yuki M","year":"2019","unstructured":"Yuki M Asano, Christian Rupprecht, and Andrea Vedaldi. 2019. A critical analysis of self-supervision, or what we can learn from a single image. arXiv preprint arXiv:1904.13132 (2019)."},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/FG47880.2020.00115"},{"key":"e_1_2_2_7_1","volume-title":"Bringing portraits to life. ACM transactions on graphics (TOG) 36, 6","author":"Averbuch-Elor Hadar","year":"2017","unstructured":"Hadar Averbuch-Elor, Daniel Cohen-Or, Johannes Kopf, and Michael F Cohen. 2017. Bringing portraits to life. ACM transactions on graphics (TOG) 36, 6 (2017), 1--13."},{"key":"e_1_2_2_8_1","volume-title":"Learning Audio-Driven Viseme Dynamics for 3D Face Animation. arXiv preprint arXiv:2301.06059","author":"Bao Linchao","year":"2023","unstructured":"Linchao Bao, Haoxian Zhang, Yue Qian, Tangli Xue, Changhai Chen, Xuefei Zhe, and Di Kang. 2023. Learning Audio-Driven Viseme Dynamics for 3D Face Animation. arXiv preprint arXiv:2301.06059 (2023)."},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964970"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459829"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1276377.1276419"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/311535.311556"},{"key":"e_1_2_2_13_1","first-page":"1","article-title":"Realistic Human Face Rendering for \"The Matrix Reloaded\". In ACM SIGGRAPH 2003 Sketches & Applications (SIGGRAPH '03). ACM, New York","volume":"16","author":"Borshukov George","year":"2003","unstructured":"George Borshukov and J. P. Lewis. 2003. Realistic Human Face Rendering for \"The Matrix Reloaded\". In ACM SIGGRAPH 2003 Sketches & Applications (SIGGRAPH '03). ACM, New York, NY, USA, 16:1--16:1.","journal-title":"NY, USA"},{"key":"e_1_2_2_14_1","volume-title":"Recognising subtle emotional expressions: The role of facial movements. Cognition and emotion 22, 8","author":"Bould Emma","year":"2008","unstructured":"Emma Bould, Neil Morris, and Brian Wink. 2008. Recognising subtle emotional expressions: The role of facial movements. Cognition and emotion 22, 8 (2008), 1569--1587."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1778765.1778778"},{"key":"e_1_2_2_16_1","first-page":"1","article-title":"Authentic volumetric avatars from a phone scan","volume":"41","author":"Cao Chen","year":"2022","unstructured":"Chen Cao, Tomas Simon, Jin Kyu Kim, Gabe Schwartz, Michael Zollhoefer, Shun-Suke Saito, Stephen Lombardi, Shih-En Wei, Danielle Belko, Shoou-I Yu, Yaser Sheikh, and Jason Saragih. 2022. Authentic volumetric avatars from a phone scan. ACM Transactions on Graphics (TOG) 41, 4 (2022), 1--19.","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925873"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01286"},{"key":"e_1_2_2_19_1","volume-title":"3D face reconstruction and gaze tracking in the HMD for virtual interaction","author":"Chen Shu-Yu","year":"2022","unstructured":"Shu-Yu Chen, Yu-Kun Lai, Shihong Xia, Paul Rosin, and Lin Gao. 2022. 3D face reconstruction and gaze tracking in the HMD for virtual interaction. IEEE Transactions on Multimedia (2022)."},{"key":"e_1_2_2_20_1","volume-title":"International conference on machine learning. PMLR, 1597--1607","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning. PMLR, 1597--1607."},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3494977"},{"key":"e_1_2_2_22_1","doi-asserted-by":"crossref","unstructured":"T. F. Cootes G. J. Edwards and C. J. Taylor. 1998. Active appearance models. In Computer Vision --- ECCV'98.","DOI":"10.1007\/BFb0054760"},{"key":"e_1_2_2_23_1","volume-title":"International conference on machine learning. PMLR, 933--941","author":"Dauphin Yann N","year":"2017","unstructured":"Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017. Language modeling with gated convolutional networks. In International conference on machine learning. PMLR, 933--941."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_2_25_1","volume-title":"Bert: Pretraining of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pretraining of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_2_2_26_1","unstructured":"Digital Domain. 2019. Project DIGI DOUG. https:\/\/digitaldomain.com\/technology\/project-digi-doug\/."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.167"},{"key":"e_1_2_2_28_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206868"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2638549"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2070781.2024163"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01810"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6717"},{"key":"e_1_2_2_35_1","volume-title":"Improved training of wasserstein gans. Advances in neural information processing systems 30","author":"Gulrajani Ishaan","year":"2017","unstructured":"Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville. 2017. Improved training of wasserstein gans. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_36_1","first-page":"5679","article-title":"Self-supervised co-training for video representation learning","volume":"33","author":"Han Tengda","year":"2020","unstructured":"Tengda Han, Weidi Xie, and Andrew Zisserman. 2020. Self-supervised co-training for video representation learning. Advances in Neural Information Processing Systems 33 (2020), 5679--5690.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00403"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_2_2_40_1","unstructured":"HTC. 2021. HTC VIVE Facial Tracker. https:\/\/www.vive.com\/eu\/accessory\/facial-tracker\/."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/1179352.1142003"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766974"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073659"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.3390\/technologies9010002"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.2352\/EI.2022.34.8.IMAGE-255"},{"key":"e_1_2_2_46_1","volume-title":"Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410","author":"Jozefowicz Rafal","year":"2016","unstructured":"Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016. Exploring the limits of language modeling. arXiv preprint arXiv:1602.02410 (2016)."},{"key":"e_1_2_2_47_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)."},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DIMPVT.2012.67"},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/2822013.2822035"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.1910338"},{"key":"e_1_2_2_51_1","volume-title":"Proceedings, Part IV 14","author":"Larsson Gustav","year":"2016","unstructured":"Gustav Larsson, Michael Maire, and Gregory Shakhnarovich. 2016. Learning representations for automatic colorization. In Computer Vision-ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part IV 14. Springer, 577--593."},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766939"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0921-8890(99)00103-7"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01252-6_6"},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201401"},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459863"},{"key":"e_1_2_2_57_1","volume-title":"Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983","author":"Loshchilov Ilya","year":"2016","unstructured":"Ilya Loshchilov and Frank Hutter. 2016. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983 (2016)."},{"key":"e_1_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00013"},{"key":"e_1_2_2_59_1","volume-title":"Meta Quest: Movement SDK for Unity - Overview and Setup. https:\/\/developer.oculus.com\/documentation\/unity\/move-overview\/#face-tracking.","author":"Meta Inc.","year":"2023","unstructured":"Meta Inc. 2023a. Meta Quest: Movement SDK for Unity - Overview and Setup. https:\/\/developer.oculus.com\/documentation\/unity\/move-overview\/#face-tracking."},{"key":"e_1_2_2_60_1","unstructured":"Meta Inc. 2023b. Meta Quest Pro: Premium Mixed Reality. https:\/\/www.meta.com\/ie\/quest\/quest-pro\/."},{"key":"e_1_2_2_61_1","volume-title":"Proceedings of the 18th International Conference on Artificial Neural Networks (ICANN), Part I (Prague, Czech Republic). 971--981","author":"Nair Vinod","unstructured":"Vinod Nair, Josh Susskind, and Geoffrey E. Hinton. 2008. Analysis-by-Synthesis by Learning to Invert Generative Black Boxes. In Proceedings of the 18th International Conference on Artificial Neural Networks (ICANN), Part I (Prague, Czech Republic). 971--981."},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46466-4_5"},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/2980179.2980252"},{"key":"e_1_2_2_64_1","volume-title":"Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32 (2019)."},{"key":"e_1_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.278"},{"key":"e_1_2_2_66_1","unstructured":"F. Pighin and J.P. Lewis. 2006. Performance-Driven Facial Animation. In ACM SIGGRAPH Courses."},{"key":"e_1_2_2_67_1","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition. 652--660","author":"Qi Charles R","year":"2017","unstructured":"Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. 2017. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition. 652--660."},{"key":"e_1_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528233.3530740"},{"key":"e_1_2_2_69_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.589"},{"key":"e_1_2_2_70_1","volume-title":"U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention-MICCAI 2015: 18th International Conference","author":"Ronneberger Olaf","year":"2015","unstructured":"Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention-MICCAI 2015: 18th International Conference, Munich, Germany, October 5--9, 2015, Proceedings, Part III 18. Springer, 234--241."},{"key":"e_1_2_2_71_1","volume-title":"Neural Strands: Learning Hair Geometry and Appearance from Multi-View Images. ECCV","author":"Rosu Radu Alexandru","year":"2022","unstructured":"Radu Alexandru Rosu, Shunsuke Saito, Ziyan Wang, Chenglei Wu, Sven Behnke, and Giljoo Nam. 2022. Neural Strands: Learning Hair Geometry and Appearance from Multi-View Images. ECCV (2022)."},{"key":"e_1_2_2_72_1","doi-asserted-by":"crossref","unstructured":"Shunsuke Saito Gabriel Schwartz Tomas Simon Junxuan Li and Giljoo Nam. 2023. Relightable Gaussian Codec Avatars. (2023). arXiv:2312.03704 [cs.GR]","DOI":"10.1109\/CVPR52733.2024.00021"},{"key":"e_1_2_2_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392493"},{"key":"e_1_2_2_74_1","doi-asserted-by":"publisher","DOI":"10.1145\/3089269.3089276"},{"key":"e_1_2_2_75_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP40776.2020.9053415"},{"key":"e_1_2_2_76_1","volume-title":"First order motion model for image animation. Advances in neural information processing systems 32","author":"Siarohin Aliaksandr","year":"2019","unstructured":"Aliaksandr Siarohin, St\u00e9phane Lathuili\u00e8re, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019. First order motion model for image animation. Advances in neural information processing systems 32 (2019)."},{"key":"e_1_2_2_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240612"},{"key":"e_1_2_2_78_1","doi-asserted-by":"publisher","DOI":"10.21437\/ICSLP.1994-538"},{"key":"e_1_2_2_79_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00270"},{"key":"e_1_2_2_80_1","volume-title":"Proceedings of the IEEE international conference on computer vision workshops. 1274--1283","author":"Tewari Ayush","year":"2017","unstructured":"Ayush Tewari, Michael Zollhofer, Hyeongwoo Kim, Pablo Garrido, Florian Bernard, Patrick Perez, and Christian Theobalt. 2017. Mofa: Model-based deep convolutional face autoencoder for unsupervised monocular reconstruction. In Proceedings of the IEEE international conference on computer vision workshops. 1274--1283."},{"key":"e_1_2_2_81_1","doi-asserted-by":"publisher","DOI":"10.1145\/2929464.2929475"},{"key":"e_1_2_2_82_1","doi-asserted-by":"publisher","DOI":"10.1145\/3182644"},{"key":"e_1_2_2_83_1","volume-title":"International conference on machine learning. PMLR, 10347--10357","author":"Touvron Hugo","year":"2021","unstructured":"Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv\u00e9 J\u00e9gou. 2021. Training data-efficient image transformers & distillation through attention. In International conference on machine learning. PMLR, 10347--10357."},{"key":"e_1_2_2_84_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00767"},{"key":"e_1_2_2_85_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_86_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01261-8_24"},{"key":"e_1_2_2_87_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3323030"},{"key":"e_1_2_2_88_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01304"},{"key":"e_1_2_2_89_1","volume-title":"Towards Practical Capture of High-Fidelity Relightable Avatars. In SIGGRAPH Asia 2023 Conference Proceedings.","author":"Yang Haotian","year":"2023","unstructured":"Haotian Yang, Mingwu Zheng, Wanquan Feng, Haibin Huang, Yu-Kun Lai, Pengfei Wan, Zhongyuan Wang, and Chongyang Ma. 2023. Towards Practical Capture of High-Fidelity Relightable Avatars. In SIGGRAPH Asia 2023 Conference Proceedings."},{"key":"e_1_2_2_90_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00457"},{"key":"e_1_2_2_91_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00612"},{"key":"e_1_2_2_92_1","volume-title":"International Conference on Machine Learning. PMLR, 12310--12320","author":"Zbontar Jure","year":"2021","unstructured":"Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and St\u00e9phane Deny. 2021. Barlow twins: Self-supervised learning via redundancy reduction. In International Conference on Machine Learning. PMLR, 12310--12320."},{"key":"e_1_2_2_93_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015706.1015759"},{"key":"e_1_2_2_94_1","volume-title":"Proceedings, Part III 14","author":"Zhang Richard","year":"2016","unstructured":"Richard Zhang, Phillip Isola, and Alexei A Efros. 2016. Colorful image colorization. In Computer Vision-ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part III 14. Springer, 649--666."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3658234","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3658234","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:04:16Z","timestamp":1750291456000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3658234"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,19]]},"references-count":93,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2024,7,19]]}},"alternative-id":["10.1145\/3658234"],"URL":"https:\/\/doi.org\/10.1145\/3658234","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,19]]},"assertion":[{"value":"2024-07-19","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}