{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T17:01:19Z","timestamp":1777568479162,"version":"3.51.4"},"reference-count":100,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T00:00:00Z","timestamp":1701734400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2023,12,5]]},"abstract":"<jats:p>\n            Immersive user experiences in live VR\/AR performances require a fast and accurate free-view rendering of the performers. Existing methods are mainly based on Pixel-aligned Implicit Functions (PIFu) or Neural Radiance Fields (NeRF). However, while PIFu-based methods usually fail to produce photorealistic view-dependent textures, NeRF-based methods typically lack local geometry accuracy and are computationally heavy (\n            <jats:italic toggle=\"yes\">e.g.<\/jats:italic>\n            , dense sampling of 3D points, additional fine-tuning, or pose estimation). In this work, we propose a novel generalizable method, named SAILOR, to create high-quality human free-view videos from very sparse RGBD live streams. To produce view-dependent textures while preserving locally accurate geometry, we integrate PIFu and NeRF such that they work synergistically by conditioning the PIFu on depth and then rendering view-dependent textures through NeRF. Specifically, we propose a novel network, named SRONet, for this hybrid representation. SRONet can handle unseen performers without fine-tuning. Besides, a neural blending-based ray interpolation approach, a tree-based voxel-denoising scheme, and a parallel computing pipeline are incorporated to reconstruct and render live free-view videos at 10 fps on average. To evaluate the rendering performance, we construct a real-captured RGBD benchmark from 40 performers. Experimental results show that SAILOR outperforms existing human reconstruction and performance capture methods.\n          <\/jats:p>","DOI":"10.1145\/3618370","type":"journal-article","created":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T10:20:48Z","timestamp":1701771648000},"page":"1-15","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["SAILOR: Synergizing Radiance and Occupancy Fields for Live Human Performance Capture"],"prefix":"10.1145","volume":"42","author":[{"given":"Zheng","family":"Dong","sequence":"first","affiliation":[{"name":"State Key Laboratory of CAD&amp;CG, Zhejiang University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ke","family":"Xu","sequence":"additional","affiliation":[{"name":"City University of Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yaoan","family":"Gao","sequence":"additional","affiliation":[{"name":"State Key Laboratory of CAD&amp;CG, Zhejiang University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qilin","family":"Sun","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen and Point Spread Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hujun","family":"Bao","sequence":"additional","affiliation":[{"name":"State Key Laboratory of CAD&amp;CG, Zhejiang University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weiwei","family":"Xu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of CAD&amp;CG, Zhejiang University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rynson W. H.","family":"Lau","sequence":"additional","affiliation":[{"name":"City University of Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,12,5]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00238"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58536-5_19"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-25082-8_3"},{"key":"e_1_2_2_4_1","volume-title":"JIFF: Jointly-aligned Implicit Face Function for High Quality Single View Clothed Human Reconstruction. In IEEE Conf. Comput. Vis. Pattern Recog.","author":"Cao Yukang","year":"2022","unstructured":"Yukang Cao, Guanying Chen, Kai Han, Wenqi Yang, and Kwan-Yee K Wong. 2022. JIFF: Jointly-aligned Implicit Face Function for High Quality Single View Clothed Human Reconstruction. In IEEE Conf. Comput. Vis. Pattern Recog."},{"key":"e_1_2_2_5_1","unstructured":"Kennard Chan Guosheng Lin Haiyu Zhao and Weisi Lin. 2022a. S-PIFu: Integrating Parametric Human Models with PIFu for Single-view Clothed Human Reconstruction. In Adv. Neural Inform. Process. Syst."},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20086-1_19"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01386"},{"key":"e_1_2_2_8_1","unstructured":"Jianchuan Chen Ying Zhang Di Kang Xuefei Zhe Linchao Bao Xu Jia and Huchuan Lu. 2021b. Animatable Neural Radiance Fields from Monocular RGB Videos. arXiv:2106.13629"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/166117.166153"},{"key":"e_1_2_2_10_1","doi-asserted-by":"crossref","unstructured":"Alvaro Collet Ming Chuang Pat Sweeney Don Gillett Dennis Evseev David Calabrese Hugues Hoppe Adam Kirk and Steve Sullivan. 2015. High-Quality Streamable Free-Viewpoint Video. ACM Trans. Graph. (2015).","DOI":"10.1145\/2766945"},{"key":"e_1_2_2_11_1","doi-asserted-by":"crossref","unstructured":"Edilson De Aguiar Carsten Stoll Christian Theobalt Naveed Ahmed Hans-Peter Seidel and Sebastian Thrun. 2008. Performance capture from sparse multi-view video. In ACM SIGGRAPH.","DOI":"10.1145\/1399504.1360697"},{"key":"e_1_2_2_12_1","doi-asserted-by":"crossref","unstructured":"Paul Debevec Yizhou Yu and George Borshukov. 1998. Efficient view-dependent image-based rendering with projective texture-mapping. In Eurographics.","DOI":"10.1007\/978-3-7091-6453-2_10"},{"key":"e_1_2_2_13_1","unstructured":"Zheng Dong Ke Xu Ziheng Duan Hujun Bao Weiwei Xu and Rynson Lau. 2022. Geometry-aware Two-scale PIFu Representation for Human Reconstruction. In Adv. Neural Inform. Process. Syst."},{"key":"e_1_2_2_14_1","unstructured":"Mingsong Dou Philip L. Davidson S. Fanello S. Khamis Adarsh Kowdle Christoph Rhemann Vladimir Tankovich and Shahram Izadi. 2017. Motion2fusion: real-time volumetric performance capture. ACM Trans. Graph. (2017)."},{"key":"e_1_2_2_15_1","volume-title":"Adarsh Kowdle, Sergio Orts Escolano, Christoph Rhemann, David Kim, Jonathan Taylor, et al.","author":"Dou Mingsong","year":"2016","unstructured":"Mingsong Dou, Sameh Khamis, Yury Degtyarev, Philip Davidson, Sean Ryan Fanello, Adarsh Kowdle, Sergio Orts Escolano, Christoph Rhemann, David Kim, Jonathan Taylor, et al. 2016. Fusion4d: Real-time performance capture of challenging scenes. ACM Trans. Graph. (2016)."},{"key":"e_1_2_2_16_1","doi-asserted-by":"crossref","unstructured":"P\u00e9ter Fankhauser Michael Bloesch Diego Rodriguez Ralf Kaestner Marco Hutter and Roland Siegwart. 2015. Kinect v2 for mobile robot navigation: Evaluation and modeling. In ICAR.","DOI":"10.1109\/ICAR.2015.7251485"},{"key":"e_1_2_2_17_1","volume-title":"FOF: Learning Fourier Occupancy Field for Monocular Real-time Human Reconstruction. In Adv. Neural Inform. Process. Syst.","author":"Feng Qiao","year":"2022","unstructured":"Qiao Feng, Yebin Liu, Yu-Kun Lai, Jingyu Yang, and Kun Li. 2022. FOF: Learning Fourier Occupancy Field for Monocular Real-time Human Reconstruction. In Adv. Neural Inform. Process. Syst."},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00854"},{"key":"e_1_2_2_19_1","volume-title":"Jongyoo Kim, Sida Peng, Zicheng Liu, and Xin Tong.","author":"Gao Xiangjun","year":"2022","unstructured":"Xiangjun Gao, Jiao long Yang, Jongyoo Kim, Sida Peng, Zicheng Liu, and Xin Tong. 2022. MPS-NeRF: Generalizable 3D Human Rendering from Multiview Images. IEEE Trans. Pattern Anal. Mach. Intell. (2022)."},{"key":"e_1_2_2_20_1","doi-asserted-by":"crossref","unstructured":"Kaiwen Guo Peter Lincoln Philip Davidson Jay Busch Xueming Yu Matt Whalen Geoff Harvey Sergio Orts-Escolano Rohit Pandey Jason Dourgarian et al. 2019. The relightables: Volumetric performance capture of humans with realistic relighting. ACM Trans. Graph. (2019).","DOI":"10.1145\/3355089.3356571"},{"key":"e_1_2_2_21_1","doi-asserted-by":"crossref","unstructured":"Marc Habermann Lingjie Liu Weipeng Xu Michael Zollhoefer Gerard Pons-Moll and Christian Theobalt. 2021. Real-time deep dynamic characters. ACM Trans. Graph. (2021).","DOI":"10.1145\/3450626.3459749"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3311970"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00510"},{"key":"e_1_2_2_24_1","doi-asserted-by":"crossref","unstructured":"Peter Hedman Julien Philip True Price Jan-Michael Frahm George Drettakis and Gabriel Brostow. 2018. Deep blending for free-viewpoint image-based rendering. ACM Trans. Graph. (2018).","DOI":"10.1145\/3272127.3275084"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00060"},{"key":"e_1_2_2_26_1","volume-title":"SHERF: Generalizable Human NeRF from a Single Image. arXiv:2303.12791","author":"Hu Shoukang","year":"2023","unstructured":"Shoukang Hu, Fangzhou Hong, Liang Pan, Haiyi Mei, Lei Yang, and Ziwei Liu. 2023. SHERF: Generalizable Human NeRF from a Single Image. arXiv:2303.12791"},{"key":"e_1_2_2_27_1","doi-asserted-by":"crossref","unstructured":"Mustafa I\u015f\u0131k Martin R\u00fcnz Markos Georgopoulos Taras Khakhulin Jonathan Starck Lourdes Agapito and Matthias Nie\u00dfner. 2023. HumanRF: High-Fidelity Neural Radiance Fields for Humans in Motion. ACM Trans. Graph. (2023).","DOI":"10.1145\/3592415"},{"key":"e_1_2_2_28_1","volume-title":"Proc. of SIGGRAPH.","author":"Jiakai Zhang","year":"2021","unstructured":"Zhang Jiakai, Liu Xinhang, Ye Xinyi, Zhao Fuqiang, Zhang Yanshun, Wu Minye, Zhang Yingliang, Xu Lan, and Yu Jingyi. 2021. Editable Free-Viewpoint Video using a Layered Neural Representation. In Proc. of SIGGRAPH."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19824-3_24"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00868"},{"key":"e_1_2_2_31_1","unstructured":"Jaehyeok Kim Dongyoon Wee and Dan Xu. 2023. You Only Train Once: Multi-Identity Free-Viewpoint Neural Human Rendering from Monocular Videos. arXiv:2303.05835"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00530"},{"key":"e_1_2_2_33_1","unstructured":"Youngjoong Kwon Dahun Kim Duygu Ceylan and Henry Fuchs. 2021. Neural human performer: Learning generalizable radiance fields for human performance rendering. In Adv. Neural Inform. Process. Syst."},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/344779.344862"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58592-1_4"},{"key":"e_1_2_2_36_1","doi-asserted-by":"crossref","unstructured":"Yue Li Marc Habermann Bernhard Thomaszewski Stelian Coros Thabo Beeler and Christian Theobalt. 2021. Deep Physics-aware Inference of Cloth Deformation for Monocular Human Performance Capture. In 3DV.","DOI":"10.1109\/3DV53792.2021.00047"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00142"},{"key":"e_1_2_2_38_1","unstructured":"Haotong Lin Sida Peng Zhen Xu Yunzhi Yan Qing Shuai Hujun Bao and Xiaowei Zhou. 2022. Efficient Neural Radiance Fields with Learned Depth-Guided Sampling. In ACM SIGGRAPH Asia."},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00865"},{"key":"e_1_2_2_40_1","volume-title":"Tat-Seng Chua, and Christian Theobalt.","author":"Liu Lingjie","year":"2020","unstructured":"Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. 2020. Neural Sparse Voxel Fields. NeurIPS."},{"key":"e_1_2_2_41_1","unstructured":"Lingjie Liu Marc Habermann Viktor Rudnev Kripasindhu Sarkar Jiatao Gu and Christian Theobalt. 2021. Neural actor: Neural free-view synthesis of human actors with pose control. ACM Trans. Graph. (2021)."},{"key":"e_1_2_2_42_1","volume-title":"A point-cloud-based multiview stereo algorithm for free-viewpoint video","author":"Liu Yebin","year":"2009","unstructured":"Yebin Liu, Qionghai Dai, and Wenli Xu. 2009. A point-cloud-based multiview stereo algorithm for free-viewpoint video. IEEE Trans. Vis. Comput. Graph. (2009)."},{"key":"e_1_2_2_43_1","doi-asserted-by":"crossref","unstructured":"Stephen Lombardi Tomas Simon Gabriel Schwartz Michael Zollhoefer Yaser Sheikh and Jason Saragih. 2021. Mixture of Volumetric Primitives for Efficient Neural Rendering. ACM Trans. Graph. (2021).","DOI":"10.1145\/3450626.3459863"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818013"},{"key":"e_1_2_2_45_1","unstructured":"Ricardo Martin-Brualla Rohit Pandey Shuoran Yang Pavel Pidlypenskyi Jonathan Taylor Julien Valentin Sameh Khamis Philip Davidson Anastasia Tkach Peter Lincoln Adarsh Kowdle Christoph Rhemann Dan B Goldman Cem Keskin Steve Seitz Shahram Izadi and Sean Fanello. 2018. LookinGood: Enhancing Performance Capture with Real-Time Neural Re-Rendering. ACM Trans. Graph. (2018)."},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/344779.344951"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19784-0_11"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58452-8_24"},{"key":"e_1_2_2_49_1","doi-asserted-by":"crossref","unstructured":"Thomas M\u00fcller Fabrice Rousselle Jan Nov\u00e1k and Alexander Keller. 2021. Real-time Neural Radiance Caching for Path Tracing. ACM Trans. Graph. (2021).","DOI":"10.1145\/3450626.3459812"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298631"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298631"},{"key":"e_1_2_2_52_1","volume-title":"Kinectfusion: Real-time dense surface mapping and tracking","author":"Newcombe Richard A","year":"2011","unstructured":"Richard A Newcombe, Shahram Izadi, Otmar Hilliges, David Molyneaux, David Kim, Andrew J Davison, Pushmeet Kohi, Jamie Shotton, Steve Hodges, and Andrew Fitzgibbon. 2011. Kinectfusion: Real-time dense surface mapping and tracking. In IEEE ISMAR."},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00554"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/2984511.2984517"},{"key":"e_1_2_2_55_1","volume-title":"Nerfies: Deformable Neural Radiance Fields. In Int. Conf. Comput. Vis.","author":"Park Keunhong","year":"2021","unstructured":"Keunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Steven M. Seitz, and Ricardo Martin-Brualla. 2021a. Nerfies: Deformable Neural Radiance Fields. In Int. Conf. Comput. Vis."},{"key":"e_1_2_2_56_1","volume-title":"Seitz","author":"Park Keunhong","year":"2021","unstructured":"Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin-Brualla, and Steven M. Seitz. 2021b. HyperNeRF: A Higher-Dimensional Representation for Topologically Varying Neural Radiance Fields. ACM Trans. Graph. (2021)."},{"key":"e_1_2_2_57_1","volume":"201","author":"Pavlakos Georgios","unstructured":"Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. 2019. Expressive Body Capture: 3D Hands, Face, and Body From a Single Image. In IEEE Conf. Comput. Vis. Pattern Recog.","journal-title":"Michael J. Black."},{"key":"e_1_2_2_58_1","doi-asserted-by":"crossref","unstructured":"Bo Peng Jun Hu Jingtao Zhou Xuan Gao and Juyong Zhang. 2023. IntrinsicNGP: Intrinsic Coordinate based Hash Encoding for Human NeRF. arXiv:2302.14683","DOI":"10.1109\/TVCG.2023.3306078"},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01405"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00894"},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01018"},{"key":"e_1_2_2_62_1","volume":"201","author":"Qi Charles R.","unstructured":"Charles R. Qi, Li Yi, Hao Su, and Leonidas J. Guibas. 2017. PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space. In Adv. Neural Inform. Process. Syst.","journal-title":"Leonidas J. Guibas."},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01550"},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/3528233.3530740"},{"key":"e_1_2_2_65_1","volume-title":"IEEE Conf. Comput. Vis. Pattern Recog.","author":"Efros Eli Shechtman Richard Alexei A","year":"2018","unstructured":"Alexei A Efros Eli Shechtman Richard Zhang, Phillip Isola and Oliver Wan. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In IEEE Conf. Comput. Vis. Pattern Recog."},{"key":"e_1_2_2_66_1","volume-title":"U-net: Convolutional networks for biomedical image segmentation. In MICCAI.","author":"Ronneberger Olaf","year":"2015","unstructured":"Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-net: Convolutional networks for biomedical image segmentation. In MICCAI."},{"key":"e_1_2_2_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00239"},{"key":"e_1_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00016"},{"key":"e_1_2_2_69_1","doi-asserted-by":"crossref","unstructured":"Ruizhi Shao Liliang Chen Zerong Zheng Hongwen Zhang Yuxiang Zhang Han Huang Yandong Guo and Yebin Liu. 2022a. FloRen: Real-Time High-Quality Human Performance Rendering via Appearance Flow Using Sparse RGB Cameras. In ACM SIGGRAPH Asia.","DOI":"10.1145\/3550469.3555409"},{"key":"e_1_2_2_70_1","volume-title":"DoubleField: Bridging the Neural Surface and Radiance Fields for High-fidelity Human Reconstruction and Rendering. In IEEE Conf. Comput. Vis. Pattern Recog.","author":"Shao Ruizhi","year":"2022","unstructured":"Ruizhi Shao, Hongwen Zhang, He Zhang, Mingjia Chen, Yanpei Cao, Tao Yu, and Yebin Liu. 2022b. DoubleField: Bridging the Neural Surface and Radiance Fields for High-fidelity Human Reconstruction and Rendering. In IEEE Conf. Comput. Vis. Pattern Recog."},{"key":"e_1_2_2_71_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19824-3_41"},{"key":"e_1_2_2_72_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20086-1_7"},{"key":"e_1_2_2_73_1","volume-title":"A-nerf: Articulated neural radiance fields for learning human shape, appearance, and pose. Adv. Neural Inform. Process. Syst.","author":"Su Shih-Yang","year":"2021","unstructured":"Shih-Yang Su, Frank Yu, Michael Zollh\u00f6fer, and Helge Rhodin. 2021. A-nerf: Articulated neural radiance fields for learning human shape, appearance, and pose. Adv. Neural Inform. Process. Syst. (2021)."},{"key":"e_1_2_2_74_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58548-8_15"},{"key":"e_1_2_2_75_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01272"},{"key":"e_1_2_2_76_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Adv. Neural Inform. Process. Syst."},{"key":"e_1_2_2_77_1","doi-asserted-by":"crossref","unstructured":"Daniel Vlasic Ilya Baran Wojciech Matusik and Jovan Popovi\u0107. 2008. Articulated Mesh Animation from Multi-View Silhouettes. ACM Trans. Graph. (2008).","DOI":"10.1145\/1399504.1360696"},{"key":"e_1_2_2_78_1","doi-asserted-by":"crossref","unstructured":"Daniel Vlasic Pieter Peers Ilya Baran Paul Debevec Jovan Popovi\u0107 Szymon Rusinkiewicz and Wojciech Matusik. 2009. Dynamic shape capture using multi-view photometric stereo. In ACM SIGGRAPH Asia.","DOI":"10.1145\/1661412.1618520"},{"key":"e_1_2_2_79_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_54"},{"key":"e_1_2_2_80_1","volume-title":"IButter: Neural Interactive Bullet Time Generator for Human Free-Viewpoint Rendering. In ACM Int. Conf. Multimedia.","author":"Wang Liao","year":"2021","unstructured":"Liao Wang, Ziyu Wang, Pei Lin, Yuheng Jiang, Xin Suo, Minye Wu, Lan Xu, and Jingyi Yu. 2021b. IButter: Neural Interactive Bullet Time Generator for Human Free-Viewpoint Rendering. In ACM Int. Conf. Multimedia."},{"key":"e_1_2_2_81_1","volume-title":"Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-Time. In IEEE Conf. Comput. Vis. Pattern Recog.","author":"Wang Liao","year":"2022","unstructured":"Liao Wang, Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yanshun Zhang, Yingliang Zhang, Minye Wu, Jingyi Yu, and Lan Xu. 2022. Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-Time. In IEEE Conf. Comput. Vis. Pattern Recog."},{"key":"e_1_2_2_82_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00466"},{"key":"e_1_2_2_83_1","volume-title":"HumanNeRF: Free-Viewpoint Rendering of Moving People From Monocular Video. In IEEE Conf. Comput. Vis. Pattern Recog.","author":"Weng Chung-Yi","year":"2022","unstructured":"Chung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron, and Ira Kemelmacher-Shlizerman. 2022. HumanNeRF: Free-Viewpoint Rendering of Moving People From Monocular Video. In IEEE Conf. Comput. Vis. Pattern Recog."},{"key":"e_1_2_2_84_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00175"},{"key":"e_1_2_2_85_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00930"},{"key":"e_1_2_2_86_1","volume":"202","author":"Xiang D.","unstructured":"D. Xiang, F. Prada, C. Wu, and J. Hodgins. 2020. MonoClothCap: Towards Temporally Coherent Clothing Capture from Monocular RGB Video. In 3DV.","journal-title":"J. Hodgins."},{"key":"e_1_2_2_87_1","volume":"202","author":"Xiu Yuliang","unstructured":"Yuliang Xiu, Jinlong Yang, Dimitrios Tzionas, and Michael J. Black. 2022. ICON: Implicit Clothed humans Obtained from Normals. In IEEE Conf. Comput. Vis. Pattern Recog.","journal-title":"Michael J. Black."},{"key":"e_1_2_2_88_1","unstructured":"Weipeng Xu Avishek Chatterjee Michael Zollh\u00f6fer Helge Rhodin Dushyant Mehta Hans-Peter Seidel and Christian Theobalt. 2018. MonoPerfCap: Human Performance Capture From Monocular Video. ACM Trans. Graph. (2018)."},{"key":"e_1_2_2_89_1","doi-asserted-by":"crossref","unstructured":"Wang Yifan Felice Serena Shihao Wu Cengiz \u00d6ztireli and Olga Sorkine-Hornung. 2019. Differentiable Surface Splatting for Point-Based Geometry Processing. ACM Trans. Graph. (2019).","DOI":"10.1145\/3355089.3356513"},{"key":"e_1_2_2_90_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00455"},{"key":"e_1_2_2_91_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00569"},{"key":"e_1_2_2_92_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00761"},{"key":"e_1_2_2_93_1","unstructured":"Jiakai Zhang Liao Wang Xinhang Liu Fuqiang Zhao Minzhang Li Haizhao Dai Boyuan Zhang Wei Yang Lan Xu and Jingyi Yu. 2022a. NeuVV: Neural Volumetric Videos with Immersive Rendering and Editing. arXiv:2202.06088"},{"key":"e_1_2_2_94_1","volume-title":"VirtualCube: An Immersive 3D Video Communication System","author":"Zhang Yizhong","unstructured":"Yizhong Zhang, Jiaolong Yang, Zhen Liu, Ruicheng Wang, Guojun Chen, Xin Tong, and Baining Guo. 2022b. VirtualCube: An Immersive 3D Video Communication System. In IEEE VR."},{"key":"e_1_2_2_95_1","doi-asserted-by":"crossref","unstructured":"Fuqiang Zhao Yuheng Jiang Kaixin Yao Jiakai Zhang Liao Wang Haizhao Dai Yuhui Zhong Yingliang Zhang Minye Wu Lan Xu and Jingyi Yu. 2022a. Human Performance Modeling and Rendering via Neural Animated Mesh. ACM Trans. Graph. (2022).","DOI":"10.1145\/3550454.3555451"},{"key":"e_1_2_2_96_1","volume-title":"HumanNeRF: Efficiently Generated Human Radiance Field From Sparse Inputs. In IEEE Conf. Comput. Vis. Pattern Recog.","author":"Zhao Fuqiang","year":"2022","unstructured":"Fuqiang Zhao, Wei Yang, Jiakai Zhang, Pei Lin, Yingliang Zhang, Jingyi Yu, and Lan Xu. 2022b. HumanNeRF: Efficiently Generated Human Radiance Field From Sparse Inputs. In IEEE Conf. Comput. Vis. Pattern Recog."},{"key":"e_1_2_2_97_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01544"},{"key":"e_1_2_2_98_1","volume-title":"PaMIR: Parametric Model-Conditioned Implicit Representation for Image-based Human Reconstruction","author":"Zheng Zerong","year":"2021","unstructured":"Zerong Zheng, Tao Yu, Yebin Liu, and Dai Qionghai. 2021. PaMIR: Parametric Model-Conditioned Implicit Representation for Image-based Human Reconstruction. IEEE Trans. Pattern Anal. Mach. Intell. (2021)."},{"key":"e_1_2_2_99_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00783"},{"key":"e_1_2_2_100_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3102128"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3618370","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3618370","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T10:48:48Z","timestamp":1755773328000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3618370"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,5]]},"references-count":100,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,12,5]]}},"alternative-id":["10.1145\/3618370"],"URL":"https:\/\/doi.org\/10.1145\/3618370","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,5]]},"assertion":[{"value":"2023-12-05","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}