{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T04:39:12Z","timestamp":1784522352615,"version":"3.55.0"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2019,7,12]],"date-time":"2019-07-12T00:00:00Z","timestamp":1562889600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2019,8,31]]},"abstract":"<jats:p>Modeling and rendering of dynamic scenes is challenging, as natural scenes often contain complex phenomena such as thin structures, evolving topology, translucency, scattering, occlusion, and biological motion. Mesh-based reconstruction and tracking often fail in these cases, and other approaches (e.g., light field video) typically rely on constrained viewing conditions, which limit interactivity. We circumvent these difficulties by presenting a learning-based approach to representing dynamic objects inspired by the integral projection model used in tomographic imaging. The approach is supervised directly from 2D images in a multi-view capture setting and does not require explicit reconstruction or tracking of the object. Our method has two primary components: an encoder-decoder network that transforms input images into a 3D volume representation, and a differentiable ray-marching operation that enables end-to-end training. By virtue of its 3D representation, our construction extrapolates better to novel viewpoints compared to screen-space rendering techniques. The encoder-decoder architecture learns a latent representation of a dynamic scene that enables us to produce novel content sequences not seen during training. To overcome memory limitations of voxel-based representations, we learn a dynamic irregular grid structure implemented with a warp field during ray-marching. This structure greatly improves the apparent resolution and reduces grid-like artifacts and jagged motion. Finally, we demonstrate how to incorporate surface-based representations into our volumetric-learning framework for applications where the highest resolution is required, using facial performance capture as a case in point.<\/jats:p>","DOI":"10.1145\/3306346.3323020","type":"journal-article","created":{"date-parts":[[2019,7,12]],"date-time":"2019-07-12T19:04:08Z","timestamp":1562958248000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":686,"title":["Neural volumes"],"prefix":"10.1145","volume":"38","author":[{"given":"Stephen","family":"Lombardi","sequence":"first","affiliation":[{"name":"Facebook Reality Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tomas","family":"Simon","sequence":"additional","affiliation":[{"name":"Facebook Reality Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jason","family":"Saragih","sequence":"additional","affiliation":[{"name":"Facebook Reality Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gabriel","family":"Schwartz","sequence":"additional","affiliation":[{"name":"Facebook Reality Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andreas","family":"Lehrmann","sequence":"additional","affiliation":[{"name":"Facebook Reality Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yaser","family":"Sheikh","sequence":"additional","affiliation":[{"name":"Facebook Reality Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,7,12]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0902-9"},{"key":"e_1_2_2_2_1","unstructured":"Agisoft. 2019. Metashape. https:\/\/www.agisoft.com\/.  Agisoft. 2019. Metashape. https:\/\/www.agisoft.com\/."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1409060.1409085"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2010324.1964970"},{"key":"e_1_2_2_5_1","volume-title":"De Bonet and Paul A. Viola","author":"Jeremy","year":"1999"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2001.937544"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/383259.383309"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766945"},{"key":"e_1_2_2_9_1","volume-title":"Deformable Convolutional Networks. In International Conference on Computer Vision (ICCV).","author":"Dai J."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8659.2012.03009.x"},{"key":"e_1_2_2_11_1","volume-title":"Image-based Rendering Using Image-based Priors. In International Conference on Computer Vision (ICCV).","author":"Fitzgibbon Andrew","year":"2005"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1561\/0600000052"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2009.161"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.13127"},{"key":"e_1_2_2_15_1","volume-title":"Multi-View Stereo for Community Photo Collections. In International Conference on Computer Vision (ICCV).","author":"Goesele Michael"},{"key":"e_1_2_2_16_1","volume-title":"Deltille Grids for Geometric Camera Calibration. In International Conference on Computer Vision (ICCV).","author":"Ha H."},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1073204.1073266"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275084"},{"key":"e_1_2_2_19_1","volume-title":"International Conference on Learning Representations (ICLR).","author":"Higgins Irina","year":"2017"},{"key":"e_1_2_2_20_1","unstructured":"Milan Ikits Joe Kniss Aaron Lefohn and Charles Hansen. 2004. GPU Gems: Programming Techniques Tips and Tricks for Real-Time Graphics (Chapter 39 Volume Rendering Techniques). Addison Wesley.  Milan Ikits Joe Kniss Aaron Lefohn and Charles Hansen. 2004. GPU Gems: Programming Techniques Tips and Tricks for Real-Time Graphics (Chapter 39 Volume Rendering Techniques). Addison Wesley."},{"key":"e_1_2_2_21_1","volume-title":"VolumeDeform: Real-Time Volumetric Non-rigid Reconstruction. In European Conference on Computer Vision (ECCV).","author":"Innmann Matthias","year":"2016"},{"key":"e_1_2_2_22_1","volume-title":"Image-to-Image Translation with Conditional Adversarial Networks. Computer Vision and Pattern Recognition (CVPR)","author":"Isola Phillip","year":"2017"},{"key":"e_1_2_2_23_1","unstructured":"Max Jaderberg Karen Simonyan Andrew Zisserman and Koray Kavukcuoglu. 2015. Spatial Transformer Networks. In Advances in Neural Information Processing Systems (NeurIPS).   Max Jaderberg Karen Simonyan Andrew Zisserman and Koray Kavukcuoglu. 2015. Spatial Transformer Networks. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2980179.2980251"},{"key":"e_1_2_2_25_1","volume-title":"International Conference on Learning Representations (ICLR).","author":"Karras Tero","year":"2018"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201283"},{"key":"e_1_2_2_27_1","volume-title":"Adam: A Method for Stochastic Optimization. In International Conference for Learning Representations (ICLR).","author":"Diederik"},{"key":"e_1_2_2_28_1","volume-title":"Auto-Encoding Variational Bayes. In International Conference on Learning Representations (ICLR).","author":"Diederik"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1008191222954"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/38.511"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/344779.344862"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201401"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275099"},{"key":"e_1_2_2_34_1","volume-title":"Real-Time Visibility-Based Fusion of Depth Maps. In International Conference on Computer Vision (ICCV).","author":"Merrell Paul","year":"2007"},{"key":"e_1_2_2_35_1","volume-title":"Seitz","author":"Newcombe Richard A.","year":"2015"},{"key":"e_1_2_2_36_1","unstructured":"Thu H Nguyen-Phuoc Chuan Li Stephen Balaban and Yongliang Yang. 2018. RenderNet: A Deep Convolutional Network for Differentiable Rendering from 3D Shapes. In Advances in Neural Information Processing Systems (NeurIPS).   Thu H Nguyen-Phuoc Chuan Li Stephen Balaban and Yongliang Yang. 2018. RenderNet: A Deep Convolutional Network for Differentiable Rendering from 3D Shapes. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2508363.2508374"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275031"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00410"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3130800.3130855"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925967"},{"key":"e_1_2_2_42_1","volume-title":"Towards Real-Time Voxel Coloring. In Image Understanding Workshop.","author":"Prock Andrew"},{"key":"e_1_2_2_43_1","volume-title":"Ali Osman Ulusoy, and Andreas Geiger","author":"Riegler Gernot","year":"2017"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.290"},{"key":"e_1_2_2_45_1","doi-asserted-by":"crossref","unstructured":"Nikolay Savinov Christian H\u00e4ne Lubor Ladicky and Marc Pollefeys. 2016. Semantic 3D Reconstruction with Continuous Regularization and Ray Potentials Using a Visibility Consistency Constraint. In Computer Vision and Pattern Recognition (CVPR).  Nikolay Savinov Christian H\u00e4ne Lubor Ladicky and Marc Pollefeys. 2016. Semantic 3D Reconstruction with Continuous Regularization and Ray Potentials Using a Visibility Consistency Constraint. In Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR.2016.589"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1014573219977"},{"key":"e_1_2_2_47_1","doi-asserted-by":"crossref","unstructured":"Johannes Lutz Sch\u00f6nberger and Jan-Michael Frahm. 2016. Structure-from-Motion Revisited. In Computer Vision and Pattern Recognition (CVPR).  Johannes Lutz Sch\u00f6nberger and Jan-Michael Frahm. 2016. Structure-from-Motion Revisited. In Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR.2016.445"},{"key":"e_1_2_2_48_1","volume-title":"Pixelwise View Selection for Unstructured Multi-View Stereo. In European Conference on Computer Vision (ECCV).","author":"Sch\u00f6nberger Johannes Lutz","year":"2016"},{"key":"e_1_2_2_49_1","volume-title":"Dyer","author":"Seitz Steven M.","year":"1997"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1008176507526"},{"key":"e_1_2_2_51_1","volume-title":"Deforming Autoencoders: Unsupervised Disentangling of Shape and Appearance. In European Conference on Computer Vision (ECCV).","author":"Shu Zhixin","year":"2018"},{"key":"e_1_2_2_52_1","doi-asserted-by":"crossref","unstructured":"V. Sitzmann J. Thies F. Heide M. Nie\u00dfner G. Wetzstein and M. Zollh\u00f6fer. 2018. DeepVoxels: Learning Persistent 3D Feature Embeddings. arXiv:1812.01024 {cs.CV} (2018).  V. Sitzmann J. Thies F. Heide M. Nie\u00dfner G. Wetzstein and M. Zollh\u00f6fer. 2018. DeepVoxels: Learning Persistent 3D Feature Embeddings. arXiv:1812.01024 {cs.CV} (2018).","DOI":"10.1109\/CVPR.2019.00254"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1008192912624"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2007.70752"},{"key":"e_1_2_2_55_1","doi-asserted-by":"crossref","unstructured":"Shubham Tulsiani Alexei A. Efros and Jitendra Malik. 2018. Multi-view Consistency as Supervisory Signal for Learning Shape and Pose Prediction. In Computer Vision and Pattern Recognition (CVPR).  Shubham Tulsiani Alexei A. Efros and Jitendra Malik. 2018. Multi-view Consistency as Supervisory Signal for Learning Shape and Pose Prediction. In Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR.2018.00306"},{"key":"e_1_2_2_56_1","doi-asserted-by":"crossref","unstructured":"Shubham Tulsiani Tinghui Zhou Alexei A. Efros and Jitendra Malik. 2017. Multi-view Supervision for Single-view Reconstruction via Differentiable Ray Consistency. In Computer Vision and Pattern Recognition (CVPR).  Shubham Tulsiani Tinghui Zhou Alexei A. Efros and Jitendra Malik. 2017. Multi-view Supervision for Single-view Reconstruction via Differentiable Ray Consistency. In Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR.2017.30"},{"key":"e_1_2_2_57_1","volume-title":"Black","author":"Ulusoy Ali Osman","year":"2015"},{"key":"e_1_2_2_58_1","volume-title":"Deep Image Prior. In Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Ulyanov Dmitry","year":"2018"},{"key":"e_1_2_2_59_1","unstructured":"Ting-Chun Wang Ming-Yu Liu Jun-Yan Zhu Guilin Liu Andrew Tao Jan Kautz and Bryan Catanzaro. 2018. Video-to-Video Synthesis. In Advances in Neural Information Processing Systems (NeurIPS).   Ting-Chun Wang Ming-Yu Liu Jun-Yan Zhu Guilin Liu Andrew Tao Jan Kautz and Bryan Catanzaro. 2018. Video-to-Video Synthesis. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661229.2661284"},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2007.4408983"},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201323"},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/2601097.2601165"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3306346.3323020","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3306346.3323020","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:25:52Z","timestamp":1750206352000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3306346.3323020"}},"subtitle":["learning dynamic renderable volumes from images"],"short-title":[],"issued":{"date-parts":[[2019,7,12]]},"references-count":63,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2019,8,31]]}},"alternative-id":["10.1145\/3306346.3323020"],"URL":"https:\/\/doi.org\/10.1145\/3306346.3323020","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,7,12]]},"assertion":[{"value":"2019-07-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}