{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,12]],"date-time":"2026-07-12T02:30:35Z","timestamp":1783823435936,"version":"3.55.0"},"reference-count":81,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2021,11,9]],"date-time":"2021-11-09T00:00:00Z","timestamp":1636416000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2022,2,28]]},"abstract":"<jats:p>We present a fully automatic system that can produce high-fidelity, photo-realistic three-dimensional (3D) digital human heads with a consumer RGB-D selfie camera. The system only needs the user to take a short selfie RGB-D video while rotating his\/her head and can produce a high-quality head reconstruction in less than 30 s. Our main contribution is a new facial geometry modeling and reflectance synthesis procedure that significantly improves the state of the art. Specifically, given the input video a two-stage frame selection procedure is first employed to select a few high-quality frames for reconstruction. Then a differentiable renderer-based 3D Morphable Model (3DMM) fitting algorithm is applied to recover facial geometries from multiview RGB-D data, which takes advantages of a powerful 3DMM basis constructed with extensive data generation and perturbation. Our 3DMM has much larger expressive capacities than conventional 3DMM, allowing us to recover more accurate facial geometry using merely linear basis. For reflectance synthesis, we present a hybrid approach that combines parametric fitting and<jats:bold>Convolutional Neural Networks (CNNs)<\/jats:bold>to synthesize high-resolution albedo\/normal maps with realistic hair\/pore\/wrinkle details. Results show that our system can produce faithful 3D digital human faces with extremely realistic details. The main code and the newly constructed 3DMM basis is publicly available.<\/jats:p>","DOI":"10.1145\/3472954","type":"journal-article","created":{"date-parts":[[2021,11,9]],"date-time":"2021-11-09T21:44:03Z","timestamp":1636494243000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":63,"title":["High-Fidelity 3D Digital Human Head Creation from RGB-D Selfies"],"prefix":"10.1145","volume":"41","author":[{"given":"Linchao","family":"Bao","sequence":"first","affiliation":[{"name":"Tencent AI Lab, Shenzhen, Guangdong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiangkai","family":"Lin","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, Guangdong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yajing","family":"Chen","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, Guangdong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haoxian","family":"Zhang","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, Guangdong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sheng","family":"Wang","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, Guangdong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xuefei","family":"Zhe","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, Guangdong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Di","family":"Kang","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, Guangdong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haozhi","family":"Huang","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, Guangdong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinwei","family":"Jiang","sequence":"additional","affiliation":[{"name":"Tencent NExT Studios, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jue","family":"Wang","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, Guangdong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dong","family":"Yu","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, Guangdong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhengyou","family":"Zhang","sequence":"additional","affiliation":[{"name":"Tencent AI Lab, Shenzhen, Guangdong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,11,9]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1667239.1667251"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.1987.4767965"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1833349.1778777"},{"key":"e_1_2_1_4_1","volume-title":"https:\/\/www.bellus3d.com\/.Retrieved","author":"D.","year":"2020","unstructured":"Bellus3 D. 2020. Bellus3D. https:\/\/www.bellus3d.com\/.Retrieved September 18, 2020 from Bellus3D. 2020. Bellus3D. https:\/\/www.bellus3d.com\/.Retrieved September 18, 2020 from"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925962"},{"key":"e_1_2_1_6_1","volume-title":"Computer Graphics Forum","author":"B\u00e9rard Pascal","unstructured":"Pascal B\u00e9rard , Derek Bradley , Markus Gross , and Thabo Beeler . 2019. Practical person-specific eye rigging . In Computer Graphics Forum , Vol. 38 . Wiley Online Library , 441\u2013454. Pascal B\u00e9rard, Derek Bradley, Markus Gross, and Thabo Beeler. 2019. Practical person-specific eye rigging. In Computer Graphics Forum, Vol. 38. Wiley Online Library, 441\u2013454."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/311535.311556"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2003.1227983"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.598"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2988458.2988490"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2461976"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE43902.2021.00111"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2013.249"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925873"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818112"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3017347"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.449"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/344779.344855"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00482"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2018.09.004"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3395208"},{"key":"e_1_2_1_23_1","volume-title":"Rendering Digital Humans in Unreal Engine 4. https:\/\/docs.unrealengine.com\/en-US\/Resources\/Showcases\/DigitalHumans\/index.html.Retrieved","year":"2020","unstructured":"EpicGames. 2020. Rendering Digital Humans in Unreal Engine 4. https:\/\/docs.unrealengine.com\/en-US\/Resources\/Showcases\/DigitalHumans\/index.html.Retrieved May 20, 2020 from EpicGames. 2020. Rendering Digital Humans in Unreal Engine 4. https:\/\/docs.unrealengine.com\/en-US\/Resources\/Showcases\/DigitalHumans\/index.html.Retrieved May 20, 2020 from"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2508363.2508380"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2890493"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.265"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00125"},{"key":"e_1_2_1_28_1","volume-title":"Proc. CVPR. IEEE, 8377\u20138386","author":"Genova Kyle","unstructured":"Kyle Genova , Forrester Cole , Aaron Maschinot , Aaron Sarna , Daniel Vlasic , and William T. Freeman . 2018. Unsupervised training for 3D morphable model regression . In Proc. CVPR. IEEE, 8377\u20138386 . Kyle Genova, Forrester Cole, Aaron Maschinot, Aaron Sarna, Daniel Vlasic, and William T. Freeman. 2018. Unsupervised training for 3D morphable model regression. In Proc. CVPR. IEEE, 8377\u20138386."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2837742"},{"key":"e_1_2_1_30_1","volume-title":"Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv:1704.04861.","author":"Howard Andrew G.","year":"2017","unstructured":"Andrew G. Howard , Menglong Zhu , Bo Chen , Dmitry Kalenichenko , Weijun Wang , Tobias Weyand , Marco Andreetto , and Hartwig Adam . 2017 . Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv:1704.04861. Retrieved from https:\/\/arxiv.org\/abs\/1704.04861. Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv:1704.04861. Retrieved from https:\/\/arxiv.org\/abs\/1704.04861."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766931"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3130800.31310887"},{"key":"e_1_2_1_33_1","unstructured":"Huirong Huang Zhiyong Wu Shiyin Kang Dongyang Dai Jia Jia Tianxiao Fu Deyi Tuo Guangzhi Lei Peng Liu Dan Su Dong Yu and Helen Meng. 2020. Speaker independent and multilingual\/mixlingual speech-driven talking head generation using phonetic posteriorgrams. arXiv:2006.11610. Retrieved from https:\/\/arxiv.org\/abs\/2006.11610. Huirong Huang Zhiyong Wu Shiyin Kang Dongyang Dai Jia Jia Tianxiao Fu Deyi Tuo Guangzhi Lei Peng Liu Dan Su Dong Yu and Helen Meng. 2020. Speaker independent and multilingual\/mixlingual speech-driven talking head generation using phonetic posteriorgrams. arXiv:2006.11610. Retrieved from https:\/\/arxiv.org\/abs\/2006.11610."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766974"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.632"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.117"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.63"},{"key":"e_1_2_1_38_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba . 2014 . Adam : A method for stochastic optimization. arXiv:1412.6980. Retrieved from https:\/\/arxiv.org\/abs\/1412.6980. Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv:1412.6980. Retrieved from https:\/\/arxiv.org\/abs\/1412.6980."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00084"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-008-0152-6"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2462019"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2462026"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2017.2739743"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275075"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2508363.2508417"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISMAR.2011.6092378"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-007-0110-8"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.29.41"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/AVSS.2009.58"},{"key":"e_1_2_1_50_1","volume-title":"Retrieved","year":"2020","unstructured":"R3ds. 2020 . Wrap 3 . Retrieved May 20, 2020 from https:\/\/www.russian3dscanner.com\/. R3ds. 2020. Wrap 3. Retrieved May 20, 2020 from https:\/\/www.russian3dscanner.com\/."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/38.946629"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.589"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2005.145"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275019"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.250"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.175"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661229.2661290"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964971"},{"key":"e_1_2_1_59_1","volume-title":"Proc. ICCV","volume":"2","author":"Tewari Ayush","year":"2017","unstructured":"Ayush Tewari , Michael Zollh\u00f6fer , Hyeongwoo Kim , Pablo Garrido , Florian Bernard , Patrick P\u00e9rez , and Christian Theobalt . 2017 . Mofa: Model-based deep convolutional face autoencoder for unsupervised monocular reconstruction . In Proc. ICCV , Vol. 2 . IEEE, 5. Ayush Tewari, Michael Zollh\u00f6fer, Hyeongwoo Kim, Pablo Garrido, Florian Bernard, Patrick P\u00e9rez, and Christian Theobalt. 2017. Mofa: Model-based deep convolutional face autoencoder for unsupervised monocular reconstruction. In Proc. ICCV, Vol. 2. IEEE, 5."},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00270"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818056"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/2929464.2929475"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.163"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00414"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00767"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275098"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/2614028.2615407"},{"key":"e_1_2_1_68_1","volume-title":"Computer Graphics Forum","author":"Wang Mengjiao","unstructured":"Mengjiao Wang , Derek Bradley , Stefanos Zafeiriou , and Thabo Beeler . 2020. Facial expression synthesis using a global-local multilinear framework . In Computer Graphics Forum , Vol. 39 . Wiley Online Library , 235\u2013245. Mengjiao Wang, Derek Bradley, Stefanos Zafeiriou, and Thabo Beeler. 2020. Facial expression synthesis using a global-local multilinear framework. In Computer Graphics Forum, Vol. 39. Wiley Online Library, 235\u2013245."},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00917"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/1186822.1073267"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964972"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/2980179.2980233"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00105"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201364"},{"key":"e_1_2_1_75_1","doi-asserted-by":"crossref","unstructured":"Chengzhu Yu Heng Lu Na Hu Meng Yu Chao Weng Kun Xu Peng Liu Deyi Tuo Shiyin Kang Guangzhi Lei Dan Su and Dong Yu. 2019. DurIAN: Duration informed attention network for multimodal synthesis. In INTERSPEECH. 2027\u20132031. Chengzhu Yu Heng Lu Na Hu Meng Yu Chao Weng Kun Xu Peng Liu Deyi Tuo Shiyin Kang Guangzhi Lei Dan Su and Dong Yu. 2019. DurIAN: Duration informed attention network for multimodal synthesis. In INTERSPEECH. 2027\u20132031.","DOI":"10.21437\/Interspeech.2020-2968"},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.23"},{"key":"e_1_2_1_77_1","volume-title":"Proc. CVPR. IEEE, 787\u2013796","author":"Zhu Xiangyu","year":"2015","unstructured":"Xiangyu Zhu , Zhen Lei , Junjie Yan , Dong Yi , and Stan Z Li . 2015 . High-fidelity pose and expression normalization for face recognition in the wild . In Proc. CVPR. IEEE, 787\u2013796 . Xiangyu Zhu, Zhen Lei, Junjie Yan, Dong Yi, and Stan Z Li. 2015. High-fidelity pose and expression normalization for face recognition in the wild. In Proc. CVPR. IEEE, 787\u2013796."},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1002\/cav.405"},{"key":"e_1_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/2601097.2601165"},{"key":"e_1_2_1_80_1","volume-title":"Computer Graphics Forum","author":"Zollh\u00f6fer Michael","unstructured":"Michael Zollh\u00f6fer , Justus Thies , Pablo Garrido , Derek Bradley , Thabo Beeler , Patrick P\u00e9rez , Marc Stamminger , Matthias Nie\u00dfner , and Christian Theobalt . 2018. State of the art on monocular 3D face reconstruction, tracking, and applications . In Computer Graphics Forum , Vol. 37 . Wiley Online Library , 523\u2013550. Michael Zollh\u00f6fer, Justus Thies, Pablo Garrido, Derek Bradley, Thabo Beeler, Patrick P\u00e9rez, Marc Stamminger, Matthias Nie\u00dfner, and Christian Theobalt. 2018. State of the art on monocular 3D face reconstruction, tracking, and applications. In Computer Graphics Forum, Vol. 37. Wiley Online Library, 523\u2013550."},{"key":"e_1_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3323044"},{"key":"e_1_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201382"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3472954","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3472954","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:11:57Z","timestamp":1750191117000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3472954"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,11,9]]},"references-count":81,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,2,28]]}},"alternative-id":["10.1145\/3472954"],"URL":"https:\/\/doi.org\/10.1145\/3472954","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,11,9]]},"assertion":[{"value":"2020-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-11-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}