{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T08:15:21Z","timestamp":1783066521251,"version":"3.54.6"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T00:00:00Z","timestamp":1783036800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2026,7,3]]},"abstract":"<jats:p>We present a diffusion-based method for relighting dynamic portrait videos with photorealism and temporal consistency. Our method is fueled by a hybrid training dataset that consists of real-captured and rendered dynamic portrait videos with diverse subject appearances, facial motions, head poses, and known lighting conditions. Specifically, we construct an LED-based lighting system for realistic lighting emulation and high-speed video relighting data acquisition. By leveraging the image priors embedded in pre-trained video diffusion models, and using per-frame high dynamic range (HDR) environment map as lighting control, we train a high-performance generative model for realistic and identity-preserving dynamic portrait video relighting. In addition to the environment map control, our model uses a synthesized background image to enable control on the camera's exposure level and color tone. Our model can produce temporally consistent relit portrait video that looks realistic and harmonious under a provided new environment and faithfully preserve the subject's expression and fine facial features, including skin tone, wrinkles, and facial hair. Our model generalizes well to unseen data, in terms of the subject appearance, motion, and lighting condition. We perform extensive experiments on relighting in-the-wild videos with various environment maps and demonstrate practical applications on portrait photography. Results show that our method achieves state-of-the-art performance in photorealism, lighting harmony, and temporal consistency. Our project page: https:\/\/yufanzhang82.github.io\/PixelCube\/.<\/jats:p>","DOI":"10.1145\/3811400","type":"journal-article","created":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:05:51Z","timestamp":1783062351000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Pixel Cube: Diffusion-based Portrait Video Relighting Through Realistic Lighting Reproduction"],"prefix":"10.1145","volume":"45","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-8814-6165","authenticated-orcid":false,"given":"Yufan","family":"Zhang","sequence":"first","affiliation":[{"name":"George Mason University, Fairfax, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6032-7736","authenticated-orcid":false,"given":"Yu","family":"Ji","sequence":"additional","affiliation":[{"name":"LightThought LLC, Fairfax, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-2959-2998","authenticated-orcid":false,"given":"Ayo","family":"Ajiboye","sequence":"additional","affiliation":[{"name":"George Mason University, Fairfax, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0133-8196","authenticated-orcid":false,"given":"Rundi","family":"Wu","sequence":"additional","affiliation":[{"name":"Columbia University, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3420-6619","authenticated-orcid":false,"given":"Yu","family":"Guo","sequence":"additional","affiliation":[{"name":"George Mason University, Fairfax, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9228-1038","authenticated-orcid":false,"given":"Changxi","family":"Zheng","sequence":"additional","affiliation":[{"name":"Columbia University, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7780-7943","authenticated-orcid":false,"given":"Jinwei","family":"Ye","sequence":"additional","affiliation":[{"name":"George Mason University, Fairfax, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,3]]},"reference":[{"key":"e_1_2_2_1_1","volume-title":"Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127","author":"Blattmann Andreas","year":"2023","unstructured":"Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, Varun Jampani, and Robin Rombach. 2023. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127 (2023)."},{"key":"e_1_2_2_2_1","volume-title":"Retrieval-Augmented Diffusion Models. arXiv preprint arXiv:2204.11824","author":"Blattmann Andreas","year":"2022","unstructured":"Andreas Blattmann, Robin Rombach, Kaan Oktay, and Bj\u00f6rn Ommer. 2022. Retrieval-Augmented Diffusion Models. arXiv preprint arXiv:2204.11824 (2022)."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00595"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00043"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2006.285"},{"key":"e_1_2_2_6_1","doi-asserted-by":"crossref","unstructured":"P Debevec A Gardner C Tchou and T Hawkins. 2004a. Postproduction re-illumination of live action using time-multiplexed lighting. Institute for Creative Technologies Technical Report No. ICT TR 5 (2004).","DOI":"10.21236\/ADA459343"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/344779.344855"},{"key":"e_1_2_2_8_1","unstructured":"Paul Debevec Chris Tchou Andrew Gardner Tim Hawkins Charis Poullis Jessi Stumpfel Andrew Jones Nathaniel Yun Per Einarsson Therese Lundgren et al. 2004b. Estimating surface reflectance properties of a complex scene under captured natural illumination. ACM Trans. Graph. (2004)."},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/566654.566614"},{"key":"e_1_2_2_10_1","volume-title":"Debevec and Jitendra Malik","author":"Paul","year":"1997","unstructured":"Paul E. Debevec and Jitendra Malik. 1997. Recovering high dynamic range radiance maps from photographs. In Proceedings of ACM SIGGRAPH."},{"key":"e_1_2_2_11_1","volume-title":"RelightVid: Temporal-Consistent Diffusion Model for Video Relighting. arXiv preprint arXiv:2501.16330","author":"Fang Ye","year":"2025","unstructured":"Ye Fang, Zeyi Sun, Shangzhan Zhang, Tong Wu, Yinghao Xu, Pan Zhang, Jiaqi Wang, Gordon Wetzstein, and Dahua Lin. 2025. RelightVid: Temporal-Consistent Diffusion Model for Video Relighting. arXiv preprint arXiv:2501.16330 (2025)."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2638549"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1882262.1866163"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2024156.2024163"},{"key":"e_1_2_2_15_1","first-page":"91","article-title":"A Dual Light Stage","volume":"5","author":"Hawkins Tim","year":"2005","unstructured":"Tim Hawkins, Per Einarsson, and Paul Debevec. 2005. A Dual Light Stage. Rendering Techniques 5, 91\u201398 (2005), 2.","journal-title":"Rendering Techniques"},{"key":"e_1_2_2_16_1","volume-title":"Proceedings of the Eurographics Conference on Rendering Techniques (EGSR).","author":"Hawkins Tim","year":"2004","unstructured":"Tim Hawkins, Andreas Wenger, Chris Tchou, Andrew Gardner, Fredrik G\u00f6ransson, and Paul Debevec. 2004. Animatable facial reflectance fields. In Proceedings of the Eurographics Conference on Rendering Techniques (EGSR)."},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3680528.3687644"},{"key":"e_1_2_2_18_1","volume-title":"Fleet","author":"Ho Jonathan","year":"2022","unstructured":"Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. 2022. Video diffusion models. arXiv:2204.03458 (2022)."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00418"},{"key":"e_1_2_2_20_1","unstructured":"Jorge Jimenez Timothy Scully Nuno Barbosa Craig Donner Xenxo Alvarez Teresa Vieira Paul Matts Ver\u00f3nica Orvalho Diego Gutierrez and Tim Weyrich. 2010. A practical appearance model for dynamic facial color. ACM Trans. Graph. (2010)."},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-4481"},{"key":"e_1_2_2_22_1","unstructured":"Noah Kadner. 2021. 1899 Wraps Innovative Virtual Production. https:\/\/theasc.com\/articles\/1899-wraps-virtual-production"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.52202\/068431-1926"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00907"},{"key":"e_1_2_2_25_1","volume-title":"Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis. arXiv preprint arXiv:2505.09358","author":"Ke Bingxin","year":"2025","unstructured":"Bingxin Ke, Kevin Qu, Tianfu Wang, Nando Metzger, Shengyu Huang, Bo Li, Anton Obukhov, and Konrad Schindler. 2025. Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis. arXiv preprint arXiv:2505.09358 (2025)."},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02371"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3543664.3543681"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_15"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3680528.3687653"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02428"},{"key":"e_1_2_2_31_1","volume-title":"LightLab: Controlling Light Sources in Images with Diffusion Models. In ACM SIGGRAPH 2025 Conference Proceedings.","author":"Magar Nadav","year":"2025","unstructured":"Nadav Magar, Amir Hertz, Eric Tabellion, Yael Pritch, Alex Rav-Acha, Ariel Shamir, and Yedid Hoshen. 2025. LightLab: Controlling Light Sources in Images with Diffusion Models. In ACM SIGGRAPH 2025 Conference Proceedings."},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00518"},{"key":"e_1_2_2_33_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Mei Yiqun","unstructured":"Yiqun Mei, Yu Zeng, He Zhang, Zhixin Shu, Xuaner Zhang, Sai Bi, Jianming Zhang, HyunJoon Jung, and Vishal M. Patel. 2024. Holo-Relighting: Controllable Volumetric Portrait Relighting from a Single Image. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417814"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459872"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3469842"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.02070"},{"key":"e_1_2_2_38_1","volume-title":"Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988","author":"Poole Ben","year":"2022","unstructured":"Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2022. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988 (2022)."},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00833"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3757377.3763962"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00617"},{"key":"e_1_2_2_42_1","volume-title":"High-Resolution Image Synthesis with Latent Diffusion Models. arXiv preprint arXiv:2112.10752","author":"Rombach Robin","year":"2021","unstructured":"Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj\u00f6rn Ommer. 2021. High-Resolution Image Synthesis with Latent Diffusion Models. arXiv preprint arXiv:2112.10752 (2021)."},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00021"},{"key":"e_1_2_2_44_1","volume-title":"BecomingLit: Relightable Gaussian Avatars with Hybrid Neural Shading. In Annual Conference on Neural Information Processing Systems.","author":"Schmidt Jonathan","year":"2025","unstructured":"Jonathan Schmidt, Simon Giebenhain, and Matthias Niessner. 2025. BecomingLit: Relightable Gaussian Avatars with Hybrid Neural Shading. In Annual Conference on Neural Information Processing Systems."},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2601097.2601137"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3095816"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417821"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528233.3530751"},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-020-00280-0"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2461944"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00044"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-022-01730-5"},{"key":"e_1_2_2_53_1","unstructured":"Andreas Wenger Andrew Gardner Chris Tchou Jonas Unger Tim Hawkins and Paul Debevec. 2005. Performance relighting and reflectance transformation with time-multiplexed illumination. ACM Trans. Graph. (2005)."},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01204"},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02148"},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3641519.3657406"},{"key":"e_1_2_2_57_1","unstructured":"Yu-Ying Yeh Koki Nagano Sameh Khamis Jan Kautz Ming-Yu Liu and Ting-Chun Wang. 2022. Learning to Relight Portrait Images via a Virtual Light Stage and Synthetic-to-Real Adaptation. ACM Trans. Graph. (2022)."},{"key":"e_1_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/3641519.3657396"},{"key":"e_1_2_2_59_1","volume-title":"Proceedings of International Conference on Learning Representations.","author":"Zhang Lvmin","year":"2025","unstructured":"Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2025. Scaling In-the-Wild Training for Diffusion-based Illumination Harmonization and Editing by Imposing Consistent Light Transport. In Proceedings of International Conference on Learning Representations."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","deposited":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:16:30Z","timestamp":1783062990000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811400"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,3]]},"references-count":59,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,7,3]]}},"alternative-id":["10.1145\/3811400"],"URL":"https:\/\/doi.org\/10.1145\/3811400","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,3]]},"assertion":[{"value":"2026-01-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}