{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T20:32:40Z","timestamp":1783715560244,"version":"3.55.0"},"reference-count":46,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T00:00:00Z","timestamp":1701734400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100011033","name":"Spanish Agencia Estatal de Investigaci\u00f3n","doi-asserted-by":"crossref","award":["PID2022-141539NB-I00"],"award-info":[{"award-number":["PID2022-141539NB-I00"]}],"id":[{"id":"10.13039\/501100011033","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2023,12,5]]},"abstract":"<jats:p>In everyday photography, physical limitations of camera sensors and lenses frequently lead to a variety of degradations in captured images such as saturation or defocus blur. A common approach to overcome these limitations is to resort to image stack fusion, which involves capturing multiple images with different focal distances or exposures. For instance, to obtain an all-in-focus image, a set of multi-focus images is captured. Similarly, capturing multiple exposures allows for the reconstruction of high dynamic range. In this paper, we present a novel approach that combines neural fields with an expressive camera model to achieve a unified reconstruction of an all-in-focus high-dynamic-range image from an image stack. Our approach is composed of a set of specialized implicit neural representations tailored to address specific sub-problems along our pipeline: We use neural implicits to predict flow to overcome misalignments arising from lens breathing, depth, and all-in-focus images to account for depth of field, as well as tonemapping to deal with sensor responses and saturation - all trained using a physically inspired supervision structure with a differentiable thin lens model at its core. An important benefit of our approach is its ability to handle these tasks simultaneously or independently, providing flexible post-editing capabilities such as refocusing and exposure adjustment. By sampling the three primary factors in photography within our framework (focal distance, aperture, and exposure time), we conduct a thorough exploration to gain valuable insights into their significance and impact on overall reconstruction quality. Through extensive validation, we demonstrate that our method outperforms existing approaches in both depth-from-defocus and all-in-focus image reconstruction tasks. Moreover, our approach exhibits promising results in each of these three dimensions, showcasing its potential to enhance captured image quality and provide greater control in post-processing.<\/jats:p>","DOI":"10.1145\/3618367","type":"journal-article","created":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T10:20:48Z","timestamp":1701771648000},"page":"1-11","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["An Implicit Neural Representation for the Image Stack: Depth, All in Focus, and High Dynamic Range"],"prefix":"10.1145","volume":"42","author":[{"given":"Chao","family":"Wang","sequence":"first","affiliation":[{"name":"Max-Planck-Institut f\u00fcr Informatik, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ana","family":"Serrano","sequence":"additional","affiliation":[{"name":"Universidad de Zaragoza, I3A, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xingang","family":"Pan","sequence":"additional","affiliation":[{"name":"Max-Planck-Institut f\u00fcr Informatik, Germany &amp; Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Krzysztof","family":"Wolski","sequence":"additional","affiliation":[{"name":"Max-Planck-Institut f\u00fcr Informatik, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bin","family":"Chen","sequence":"additional","affiliation":[{"name":"Max-Planck-Institut f\u00fcr Informatik, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Karol","family":"Myszkowski","sequence":"additional","affiliation":[{"name":"Max-Planck-Institut f\u00fcr Informatik, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hans-Peter","family":"Seidel","sequence":"additional","affiliation":[{"name":"Max-Planck-Institut f\u00fcr Informatik, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christian","family":"Theobalt","sequence":"additional","affiliation":[{"name":"Max-Planck-Institut f\u00fcr Informatik, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thomas","family":"Leimk\u00fchler","sequence":"additional","affiliation":[{"name":"Max-Planck-Institut f\u00fcr Informatik, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,12,5]]},"reference":[{"key":"e_1_2_2_1_1","volume-title":"2021 Picture Coding Symposium (PCS). IEEE, 1--5.","author":"Maryam","unstructured":"Maryam Azimi et al. 2021. PU21: A novel perceptually uniform encoding for adapting existing quality metrics for HDR. In 2021 Picture Coding Symposium (PCS). IEEE, 1--5."},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459775"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2922097"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/258734.258884"},{"key":"e_1_2_2_5_1","volume-title":"Depth map prediction from a single image using a multi-scale deep network. Advances in neural information processing systems 27","author":"Eigen David","year":"2014","unstructured":"David Eigen, Christian Puhrsch, and Rob Fergus. 2014. Depth map prediction from a single image using a multi-scale deep network. Advances in neural information processing systems 27 (2014)."},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2010.5540089"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3550469.3555417"},{"key":"e_1_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Herbert Gross. 2005. Handbook of Optical Systems. (2005).","DOI":"10.1002\/9783527699223"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-24785-9_38"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2007.4408898"},{"key":"e_1_2_2_11_1","volume-title":"Maximilian Christian Staab, Laura Leal-Taix\u00e9, and Daniel Cremers.","author":"Hazirbas Caner","year":"2019","unstructured":"Caner Hazirbas, Sebastian Georg Soyer, Maximilian Christian Staab, Laura Leal-Taix\u00e9, and Daniel Cremers. 2019. Deep depth from focus. In Computer Vision-ACCV 2018: 14th Asian Conference on Computer Vision, Perth, Australia, December 2--6, 2018, Revised Selected Papers, Part III 14. Springer, 525--541."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01785"},{"key":"e_1_2_2_13_1","volume-title":"Manual of Photography","author":"Jacobson Ralph","unstructured":"Ralph Jacobson, Sidney Ray, Geoffrey G Attridge, and Norman Axford. 2000. Manual of Photography. Taylor & Francis."},{"key":"e_1_2_2_14_1","volume-title":"Proceedings, Part II 14","author":"Johnson Justin","year":"2016","unstructured":"Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016. Perceptual losses for real-time style transfer and super-resolution. In Computer Vision-ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part II 14. Springer, 694--711."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19824-3_23"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3478513.3480546"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2016.32"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.365"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2013.2244222"},{"key":"e_1_2_2_20_1","volume-title":"Computer Vision-ACCV 2012: 11th Asian Conference on Computer Vision, Daejeon, Korea, November 5--9","author":"Lin Xing","year":"2012","unstructured":"Xing Lin, Jinli Suo, Xun Cao, and Qionghai Dai. 2013. Iterative feedback estimation of depth and radiance from defocused images. In Computer Vision-ACCV 2012: 11th Asian Conference on Computer Vision, Daejeon, Korea, November 5--9, 2012, Revised Selected Papers, Part IV 11. Springer, 95--109."},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2016.12.001"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00172"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01252"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISMAR52148.2021.00068"},{"key":"e_1_2_2_25_1","volume-title":"HDR-VDP-3: A multi-metric for predicting image differences, quality and contrast distortions in high dynamic range and regular content. arXiv preprint arXiv:2304.13625","author":"Mantiuk Rafal K","year":"2023","unstructured":"Rafal K Mantiuk, Dounia Hammou, and Param Hanji. 2023. HDR-VDP-3: A multi-metric for predicting image differences, quality and contrast distortions in high dynamic range and regular content. arXiv preprint arXiv:2304.13625 (2023)."},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00115"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/PG.2007.17"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01571"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503250"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2015.2479469"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528223.3530127"},{"key":"e_1_2_2_32_1","volume-title":"Proceedings, Part VII. Springer, 216--232","author":"Nam Seonghyeon","year":"2022","unstructured":"Seonghyeon Nam, Marcus A Brubaker, and Michael S Brown. 2022. Neural image representations for multi-image fusion and layer separation. In Computer Vision-ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23--27, 2022, Proceedings, Part VII. Springer, 216--232."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/965161.806818"},{"key":"e_1_2_2_34_1","volume-title":"Fully Self-Supervised Depth Estimation from Defocus Clue. arXiv preprint arXiv:2303.10752","author":"Si Haozhe","year":"2023","unstructured":"Haozhe Si, Bin Zhao, Dong Wang, Yupeng Gao, Mulin Chen, Zhigang Wang, and Xuelong Li. 2023. Fully Self-Supervised Depth Estimation from Defocus Clue. arXiv preprint arXiv:2303.10752 (2023)."},{"key":"e_1_2_2_35_1","volume-title":"Indoor segmentation and support inference from rgbd images. ECCV (5) 7576","author":"Silberman Nathan","year":"2012","unstructured":"Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. 2012. Indoor segmentation and support inference from rgbd images. ECCV (5) 7576 (2012), 746--760."},{"key":"e_1_2_2_36_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_2_2_37_1","first-page":"7462","article-title":"Implicit neural representations with periodic activation functions","volume":"33","author":"Sitzmann Vincent","year":"2020","unstructured":"Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. 2020. Implicit neural representations with periodic activation functions. Advances in Neural Information Processing Systems 33 (2020), 7462--7473.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298972"},{"key":"e_1_2_2_39_1","volume-title":"GlowGAN: Unsupervised Learning of HDR Images from LDR Images in the Wild. arXiv preprint arXiv:2211.12352","author":"Wang Chao","year":"2022","unstructured":"Chao Wang, Ana Serrano, Xingang Pan, Bin Chen, Hans-Peter Seidel, Christian Theobalt, Karol Myszkowski, and Thomas Leimkuehler. 2022. GlowGAN: Unsupervised Learning of HDR Images from LDR Images in the Wild. arXiv preprint arXiv:2211.12352 (2022)."},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01239"},{"key":"e_1_2_2_41_1","volume-title":"Image quality assessment: from error visibility to structural similarity","author":"Wang Zhou","year":"2004","unstructured":"Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13, 4 (2004), 600--612."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.159901"},{"key":"e_1_2_2_43_1","volume-title":"Proceedings, Part I. Springer, 1--18","author":"Won Changyeon","year":"2022","unstructured":"Changyeon Won and Hae-Gon Jeon. 2022. Learning Depth from Focus in the Wild. In Computer Vision-ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23--27, 2022, Proceedings, Part I. Springer, 1--18."},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548088"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01231"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2524212"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3618367","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3618367","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T10:46:34Z","timestamp":1755773194000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3618367"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,5]]},"references-count":46,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,12,5]]}},"alternative-id":["10.1145\/3618367"],"URL":"https:\/\/doi.org\/10.1145\/3618367","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,5]]},"assertion":[{"value":"2023-12-05","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}