{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T11:44:44Z","timestamp":1783338284798,"version":"3.54.6"},"reference-count":37,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,6,5]],"date-time":"2023-06-05T00:00:00Z","timestamp":1685923200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2023,6,30]]},"abstract":"<jats:p>We propose a physically motivated deep learning framework to solve a general version of the challenging indoor lighting estimation problem. Given a single LDR image with a depth map, our method predicts spatially consistent lighting at any given image position. Particularly, when the input is an LDR video sequence, our framework not only progressively refines the lighting prediction as it sees more regions, but also preserves temporal consistency by keeping the refinement smooth. Our framework reconstructs a spherical Gaussian lighting volume (SGLV) through a tailored 3D encoder-decoder, which enables spatially consistent lighting prediction through volume ray tracing, a hybrid blending network for detailed environment maps, an in-network Monte Carlo rendering layer to enhance photorealism for virtual object insertion, and recurrent neural networks (RNN) to achieve temporally consistent lighting prediction with a video sequence as the input. For training, we significantly enhance the OpenRooms public dataset of photorealistic synthetic indoor scenes with around 360k HDR environment maps of much higher resolution and 38k video sequences, rendered with GPU-based path tracing. Experiments show that our framework achieves lighting prediction with higher quality compared to state-of-the-art single-image or video-based methods, leading to photorealistic AR applications such as object insertion.<\/jats:p>","DOI":"10.1145\/3595921","type":"journal-article","created":{"date-parts":[[2023,5,5]],"date-time":"2023-05-05T12:14:16Z","timestamp":1683288856000},"page":"1-15","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Spatiotemporally Consistent HDR Indoor Lighting Estimation"],"prefix":"10.1145","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0868-2141","authenticated-orcid":false,"given":"Zhengqin","family":"Li","sequence":"first","affiliation":[{"name":"Meta Reality Labs Research, UC San Diego, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-4420-4219","authenticated-orcid":false,"given":"Li","family":"Yu","sequence":"additional","affiliation":[{"name":"Meta Reality Labs, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9851-4445","authenticated-orcid":false,"given":"Mikhail","family":"Okunev","sequence":"additional","affiliation":[{"name":"Meta Reality Labs Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4683-2454","authenticated-orcid":false,"given":"Manmohan","family":"Chandraker","sequence":"additional","affiliation":[{"name":"UC San Diego, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9026-6886","authenticated-orcid":false,"given":"Zhao","family":"Dong","sequence":"additional","affiliation":[{"name":"Meta Reality Labs Research, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,6,5]]},"reference":[{"key":"e_1_3_1_2_1","first-page":"11441","volume-title":"Proceedings of the CVPR","author":"Akimoto Naofumi","year":"2022","unstructured":"Naofumi Akimoto, Yuhi Matsuo, and Yoshimitsu Aoki. 2022. Diverse plausible 360-degree image outpainting for efficient 3DCG background creation. In Proceedings of the CVPR. 11441\u201311450."},{"key":"e_1_3_1_3_1","volume-title":"Proceedings of the CVPR","author":"Barron Jonathan T.","year":"2013","unstructured":"Jonathan T. Barron and Jitendra Malik. 2013. Intrinsic scene properties from a single RGB-D image. In Proceedings of the CVPR."},{"key":"e_1_3_1_4_1","article-title":"Learning phrase representations using RNN encoder-decoder for statistical machine translation","author":"Cho Kyunghyun","year":"2014","unstructured":"Kyunghyun Cho, Bart Van Merri\u00ebnboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078 (2014).","journal-title":"arXiv preprint arXiv:1406.1078"},{"key":"e_1_3_1_5_1","volume-title":"Proceedings of the ECCV","author":"Choy Christopher B.","year":"2016","unstructured":"Christopher B. Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 2016. 3D-R2N2: A unified approach for single and multi-view 3D object reconstruction. In Proceedings of the ECCV."},{"key":"e_1_3_1_6_1","article-title":"Guided co-modulated GAN for 360 \\(^{\\circ }\\) field of view extrapolation","author":"Dastjerdi Mohammad Reza Karimi","year":"2022","unstructured":"Mohammad Reza Karimi Dastjerdi, Yannick Hold-Geoffroy, Jonathan Eisenmann, Siavash Khodadadeh, and Jean-Fran\u00e7ois Lalonde. 2022. Guided co-modulated GAN for 360 \\(^{\\circ }\\) field of view extrapolation. arXiv preprint arXiv:2204.07286 (2022).","journal-title":"arXiv preprint arXiv:2204.07286"},{"key":"e_1_3_1_7_1","first-page":"189","volume-title":"Proceedings of the SIGGRAPH","author":"Debevec Paul","year":"1998","unstructured":"Paul Debevec. 1998. Rendering synthetic objects into real scenes: Bridging traditional and image-based graphics with global illumination and high dynamic range photography. In Proceedings of the SIGGRAPH. 189\u2013198."},{"key":"e_1_3_1_8_1","volume-title":"Proceedings of the ICCV","author":"Gardner Marc-Andr\u00e9","year":"2019","unstructured":"Marc-Andr\u00e9 Gardner, Yannick Hold-Geoffroy, Kalyan Sunkavalli, Christian Gagn\u00e9, and Jean-Fran\u00e7ois Lalonde. 2019. Deep parametric indoor lighting estimation. In Proceedings of the ICCV."},{"issue":"4","key":"e_1_3_1_9_1","article-title":"Learning to predict indoor illumination from a single image","volume":"9","author":"Gardner Marc-Andr\u00e9","year":"2017","unstructured":"Marc-Andr\u00e9 Gardner, Kalyan Sunkavalli, Ersin Yumer, Xiaohui Shen, Emiliano Gambaretto, Christian Gagn\u00e9, and Jean-Fran\u00e7ois Lalonde. 2017. Learning to predict indoor illumination from a single image. ACM Trans. Graph. 9, 4 (2017).","journal-title":"ACM Trans. Graph."},{"key":"e_1_3_1_10_1","volume-title":"Proceedings of the CVPR","author":"Garon Mathieu","year":"2019","unstructured":"Mathieu Garon, Kalyan Sunkavalli, Sunil Hadap, Nathan Carr, and Jean-Fran\u00e7ois Lalonde. 2019. Fast spatially-varying indoor lighting estimation. In Proceedings of the CVPR."},{"key":"e_1_3_1_11_1","doi-asserted-by":"crossref","unstructured":"Brian Karis and Epic Games. 2013. Real shading in unreal engine 4. Proc. Physically Based Shading Theory Practice 4 3 (2013) 1 pages.","DOI":"10.1145\/2504435.2504457"},{"key":"e_1_3_1_12_1","doi-asserted-by":"crossref","unstructured":"Kevin Karsch Kalyan Sunkavalli Sunil Hadap Nathan Carr Hailin Jin Rafael Fonte Michael Sittig and David Forsyth. 2014. Automatic scene inference for 3d object compositing. ACM Transactions on Graphics (TOG) 33 3 (2014) 1\u201315.","DOI":"10.1145\/2602146"},{"key":"e_1_3_1_13_1","article-title":"Adam: A method for stochastic optimization","author":"Kingma Diederik","year":"2014","unstructured":"Diederik Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014).","journal-title":"arXiv preprint arXiv:1412.6980"},{"key":"e_1_3_1_14_1","first-page":"5918","volume-title":"Proceedings of the CVPR","author":"LeGendre Chloe","year":"2019","unstructured":"Chloe LeGendre, Wan-Chun Ma, Graham Fyffe, John Flynn, Laurent Charbonnel, Jay Busch, and Paul Debevec. 2019. DeepLight: Learning illumination for unconstrained mobile mixed reality. In Proceedings of the CVPR. 5918\u20135928."},{"key":"e_1_3_1_15_1","volume-title":"Proceedings of the BMVC","author":"Li Wenbin","year":"2018","unstructured":"Wenbin Li, Sajad Saeedi, John McCormac, Ronald Clark, Dimos Tzoumanikas, Qing Ye, Yuzhong Huang, Rui Tang, and Stefan Leutenegger. 2018. InteriorNet: Mega-scale multi-sensor photo-realistic indoor scenes dataset. In Proceedings of the BMVC."},{"key":"e_1_3_1_16_1","volume-title":"Proceedings of the CVPR","author":"Li Zhengqin","year":"2020","unstructured":"Zhengqin Li, Mohammad Shafiei, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. 2020. Inverse rendering for complex indoor scenes: Shape, spatially-varying lighting and SVBRDF from a single image. In Proceedings of the CVPR."},{"key":"e_1_3_1_17_1","article-title":"Physically-based editing of indoor scene lighting from a single image","author":"Li Zhengqin","year":"2022","unstructured":"Zhengqin Li, Jia Shi, Sai Bi, Rui Zhu, Kalyan Sunkavalli, Milo\u0161 Ha\u0161an, Zexiang Xu, Ravi Ramamoorthi, and Manmohan Chandraker. 2022. Physically-based editing of indoor scene lighting from a single image. arXiv preprint arXiv:2205.09343 (2022).","journal-title":"arXiv preprint arXiv:2205.09343"},{"key":"e_1_3_1_18_1","article-title":"OpenRooms: An end-to-end open framework for photorealistic indoor scene datasets","author":"Li Zhengqin","year":"2020","unstructured":"Zhengqin Li, Ting-Wei Yu, Shen Sang, Sarah Wang, Meng Song, Yuhan Liu, Yu-Ying Yeh, Rui Zhu, Nitesh Gundavarapu, Jia Shi et\u00a0al. 2020b. OpenRooms: An end-to-end open framework for photorealistic indoor scene datasets. Proceedings of the CVPR.","journal-title":"Proceedings of the CVPR"},{"key":"e_1_3_1_19_1","first-page":"10986","volume-title":"Proceedings of the CVPR","author":"Liu Chao","year":"2019","unstructured":"Chao Liu, Jinwei Gu, Kihwan Kim, Srinivasa G. Narasimhan, and Jan Kautz. 2019. Neural RGB->D sensing: Depth and uncertainty from a video camera. In Proceedings of the CVPR. 10986\u201310995."},{"issue":"4","key":"e_1_3_1_20_1","first-page":"71","article-title":"Consistent video depth estimation","volume":"39","author":"Luo Xuan","year":"2020","unstructured":"Xuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen, and Johannes Kopf. 2020. Consistent video depth estimation. ACM Trans. Graph. 39, 4 (2020), 71\u20131.","journal-title":"ACM Trans. Graph."},{"key":"e_1_3_1_21_1","first-page":"2678","volume-title":"Proceedings of the ICCV","author":"McCormac John","year":"2017","unstructured":"John McCormac, Ankur Handa, Stefan Leutenegger, and Andrew J. Davison. 2017. SceneNet RGB-D: Can 5m synthetic images beat generic ImageNet pre-training on indoor segmentation? In Proceedings of the ICCV. 2678\u20132687."},{"key":"e_1_3_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3355089.3356528"},{"key":"e_1_3_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/383259.383317"},{"key":"e_1_3_1_24_1","volume-title":"Proceedings of the ECCV","author":"Sch\u00f6nberger Johannes Lutz","year":"2016","unstructured":"Johannes Lutz Sch\u00f6nberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. 2016. Pixelwise view selection for unstructured multi-view stereo. In Proceedings of the ECCV."},{"key":"e_1_3_1_25_1","volume-title":"Proceedings of the ICCV","author":"Sengupta Soumyadip","year":"2019","unstructured":"Soumyadip Sengupta, Jinwei Gu, Kihwan Kim, Guilin Liu, David W. Jacobs, and Jan Kautz. 2019. Neural inverse rendering of an indoor scene from a single image. In Proceedings of the ICCV."},{"key":"e_1_3_1_26_1","volume-title":"Proceedings of the ECCV","author":"Silberman Nathan","year":"2012","unstructured":"Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. 2012. Indoor segmentation and support inference from RGBD images. In Proceedings of the ECCV."},{"key":"e_1_3_1_27_1","first-page":"2437","volume-title":"Proceedings of the CVPR","author":"Sitzmann Vincent","year":"2019","unstructured":"Vincent Sitzmann, Justus Thies, Felix Heide, Matthias Nie\u00dfner, Gordon Wetzstein, and Michael Zollhofer. 2019. DeepVoxels: Learning persistent 3D feature embeddings. In Proceedings of the CVPR. 2437\u20132446."},{"key":"e_1_3_1_28_1","first-page":"11298","volume-title":"Proceedings of the CVPR","author":"Somanath Gowri","year":"2021","unstructured":"Gowri Somanath and Daniel Kurz. 2021. HDR environment map estimation for real-time augmented reality. In Proceedings of the CVPR. 11298\u201311306."},{"key":"e_1_3_1_29_1","first-page":"6918","volume-title":"Proceedings of the CVPR","author":"Song Shuran","year":"2019","unstructured":"Shuran Song and Thomas Funkhouser. 2019. Neural illumination: Lighting prediction for indoor environments. In Proceedings of the CVPR. 6918\u20136926."},{"key":"e_1_3_1_30_1","first-page":"8080","volume-title":"Proceedings of the CVPR","author":"Srinivasan Pratul P.","year":"2020","unstructured":"Pratul P. Srinivasan, Ben Mildenhall, Matthew Tancik, Jonathan T. Barron, Richard Tucker, and Noah Snavely. 2020. Lighthouse: Predicting lighting volumes for spatially-coherent illumination. In Proceedings of the CVPR. 8080\u20138089."},{"key":"e_1_3_1_31_1","first-page":"12538","volume-title":"Proceedings of the ICCV","author":"Wang Zian","year":"2021","unstructured":"Zian Wang, Jonah Philion, Sanja Fidler, and Jan Kautz. 2021. Learning indoor inverse rendering with 3D spatially-varying lighting. In Proceedings of the ICCV. 12538\u201312547."},{"key":"e_1_3_1_32_1","article-title":"Object-based illumination estimation with rendering-aware neural networks","author":"Wei Xin","year":"2020","unstructured":"Xin Wei, Guojun Chen, Yue Dong, Stephen Lin, and Xin Tong. 2020. Object-based illumination estimation with rendering-aware neural networks. Proceedings of the ECCV.","journal-title":"Proceedings of the ECCV"},{"key":"e_1_3_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3151997"},{"key":"e_1_3_1_34_1","first-page":"12830","volume-title":"Proceedings of the ICCV","author":"Zhan Fangneng","year":"2021","unstructured":"Fangneng Zhan, Changgong Zhang, Wenbo Hu, Shijian Lu, Feiying Ma, Xuansong Xie, and Ling Shao. 2021a. Sparse needlets for lighting estimation with spherical transport loss. In Proceedings of the ICCV. 12830\u201312839."},{"key":"e_1_3_1_35_1","first-page":"3287","volume-title":"Proceedings of the AAAI","author":"Zhan Fangneng","year":"2021","unstructured":"Fangneng Zhan, Changgong Zhang, Yingchen Yu, Yuan Chang, Shijian Lu, Feiying Ma, and Xuansong Xie. 2021b. EMLight: Lighting estimation via spherical distribution approximation. In Proceedings of the AAAI. 3287\u20133295."},{"issue":"4","key":"e_1_3_1_36_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3450626.3459871","article-title":"Consistent depth of moving objects in video","volume":"40","author":"Zhang Zhoutong","year":"2021","unstructured":"Zhoutong Zhang, Forrester Cole, Richard Tucker, William T. Freeman, and Tali Dekel. 2021. Consistent depth of moving objects in video. ACM Trans. Graph. 40, 4 (2021), 1\u201312.","journal-title":"ACM Trans. Graph."},{"key":"e_1_3_1_37_1","first-page":"678","volume-title":"Proceedings of the ECCV","author":"Zhao Yiqin","year":"2020","unstructured":"Yiqin Zhao and Tian Guo. 2020. PointAR: Efficient lighting estimation for mobile augmented reality. In Proceedings of the ECCV. Springer, 678\u2013693."},{"key":"e_1_3_1_38_1","first-page":"2822","volume-title":"Proceedings of the CVPR","author":"Zhu Rui","year":"2022","unstructured":"Rui Zhu, Zhengqin Li, Janarbek Matai, Fatih Porikli, and Manmohan Chandraker. 2022. IRISformer: Dense vision transformers for single-image inverse rendering in indoor scenes. In Proceedings of the CVPR. 2822\u20132831."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3595921","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3595921","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:35:56Z","timestamp":1750178156000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3595921"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,5]]},"references-count":37,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,6,30]]}},"alternative-id":["10.1145\/3595921"],"URL":"https:\/\/doi.org\/10.1145\/3595921","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,5]]},"assertion":[{"value":"2022-04-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-24","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-06-05","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}