{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T09:40:56Z","timestamp":1777110056333,"version":"3.51.4"},"reference-count":103,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T00:00:00Z","timestamp":1777075200000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T00:00:00Z","timestamp":1777075200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Immersive applications call for synthesizing spatiotemporal 4D content from casual videos without costly 3D supervision. Existing video\u2010to\u20104D methods typically rely on manually annotated camera poses, which are labor\u2010intensive and brittle for in\u2010the\u2010wild footage. Recent warp\u2010then\u2010inpaint approaches mitigate the need for pose labels by warping input frames along a novel camera trajectory and using an inpainting model to fill missing regions, thereby depicting the 4D scene from diverse viewpoints. However, this trajectory\u2010to\u2010trajectory formulation often entangles camera motion with scene dynamics and complicates both modeling and inference. We introduce\n                    <jats:italic>\n                      S\n                      <jats:sc>ee<\/jats:sc>\n                      4D\n                    <\/jats:italic>\n                    , a pose\u2010free, trajectory\u2010to\u2010camera framework that replaces explicit trajectory prediction with rendering to a bank of fixed virtual cameras, thereby separating camera control from scene modeling. A view\u2010conditional video inpainting model is trained to learn a robust geometry prior by denoising realistically synthesized warped images and to inpaint occluded or missing regions across virtual viewpoints, eliminating the need for explicit 3D annotations. Building on this inpainting core, we design a spatiotemporal autoregressive inference pipeline that traverses virtual\u2010camera splines and extends videos with overlapping windows, enabling coherent generation at bounded per\u2010step complexity. We validate See4D on cross\u2010view video generation and sparse reconstruction benchmarks. Across quantitative metrics and qualitative assessments, our method achieves superior generalization and improved performance relative to pose\u2010 or trajectory\u2010conditioned baselines, advancing practical 4D world modeling from casual videos.\n                  <\/jats:p>","DOI":"10.1111\/cgf.70345","type":"journal-article","created":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T08:51:32Z","timestamp":1777107092000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["S\n                    <scp>ee<\/scp>\n                    4D: Pose\u2010Free 4D Generation via Auto\u2010Regressive Video Inpainting"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-7869-9970","authenticated-orcid":false,"given":"Dongyue","family":"Lu","sequence":"first","affiliation":[{"name":"NUS"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5969-4022","authenticated-orcid":false,"given":"Ao","family":"Liang","sequence":"additional","affiliation":[{"name":"NUS"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3579-371X","authenticated-orcid":false,"given":"Tianxin","family":"Huang","sequence":"additional","affiliation":[{"name":"HKU"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3321-8167","authenticated-orcid":false,"given":"Xiao","family":"Fu","sequence":"additional","affiliation":[{"name":"CUHK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4754-0325","authenticated-orcid":false,"given":"Yuyang","family":"Zhao","sequence":"additional","affiliation":[{"name":"NUS"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7229-2386","authenticated-orcid":false,"given":"Baorui","family":"Ma","sequence":"additional","affiliation":[{"name":"THU"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1821-4296","authenticated-orcid":false,"given":"Liang","family":"Pan","sequence":"additional","affiliation":[{"name":"Shanghai AI Lab"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4349-8297","authenticated-orcid":false,"given":"Wei","family":"Yin","sequence":"additional","affiliation":[{"name":"Horizon Robotics"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3884-2185","authenticated-orcid":false,"given":"Lingdong","family":"Kong","sequence":"additional","affiliation":[{"name":"NUS"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8994-1736","authenticated-orcid":false,"given":"Wei Tsang","family":"Ooi","sequence":"additional","affiliation":[{"name":"NUS"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4220-5958","authenticated-orcid":false,"given":"Ziwei","family":"Liu","sequence":"additional","affiliation":[{"name":"NTU"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,4,25]]},"reference":[{"key":"e_1_2_6_2_2","first-page":"1","volume-title":"2016 IEEE aerospace conference","author":"Anthes Christoph","year":"2016"},{"key":"e_1_2_6_3_2","volume-title":"Virtual reality technology","author":"Burdea Grigore C","year":"2003"},{"key":"e_1_2_6_4_2","unstructured":"Blattmann Andreas Dockhorn Tim Kulal Sumith et al. \u201cStable video diffusion: Scaling latent video diffusion models to large datasets\u201d.arXiv preprint arXiv:2311.15127(2023) 3."},{"key":"e_1_2_6_5_2","doi-asserted-by":"crossref","unstructured":"Bian Weikang Huang Zhaoyang Shi Xiaoyu et al. \u201cGS-DiT: Advancing Video Generation with Pseudo 4D Gaussian Fields through Efficient Dense 3D Point Tracking\u201d.arXiv preprint arXiv:2501.02690(2025) 3.","DOI":"10.1109\/CVPR52734.2025.02023"},{"key":"e_1_2_6_6_2","doi-asserted-by":"crossref","unstructured":"Barron Jonathan T Mildenhall Ben Tancik Matthew et al. \u201cMip-nerf: A multiscale representation for antialiasing neural radiance fields\u201d.IEEE\/CVF International Conference on Computer Vision.2021 5855\u201358643.","DOI":"10.1109\/ICCV48922.2021.00580"},{"key":"e_1_2_6_7_2","doi-asserted-by":"crossref","unstructured":"Barron Jonathan T Mildenhall Ben Verbin Dor et al. \u201cMip-nerf 360: Unbounded anti-aliased neural radiance fields\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2022 5470\u201354793.","DOI":"10.1109\/CVPR52688.2022.00539"},{"key":"e_1_2_6_8_2","doi-asserted-by":"crossref","unstructured":"Barron Jonathan T Mildenhall Ben Verbin Dor et al. \u201cZip-nerf: Anti-aliased grid-based neural radiance fields\u201d.IEEE\/CVF International Conference on Computer Vision.2023 19697\u2013197053.","DOI":"10.1109\/ICCV51070.2023.01804"},{"key":"e_1_2_6_9_2","doi-asserted-by":"crossref","unstructured":"Bain Max Nagrani Arsha Varol G\u00fcl andZisserman Andrew. \u201cFrozen in time: A joint video and image encoder for end-to-endretrieval\u201d.Proceedings of the IEEE\/CVF international conference on computer vision.2021 1728\u201317386.","DOI":"10.1109\/ICCV48922.2021.00175"},{"key":"e_1_2_6_10_2","unstructured":"Bai Jianhong Xia Menghan Fu Xiao et al. \u201cReCamMaster: Camera-Controlled Generative Rendering from A Single Video\u201d.arXiv preprint arXiv:2503.11647(2025) 2 3 6 7."},{"key":"e_1_2_6_11_2","unstructured":"Bai Jianhong Xia Menghan Wang Xintao et al. \u201cSynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints\u201d.arXiv preprint arXiv:2412.07760(2024) 2 3 6."},{"key":"e_1_2_6_12_2","doi-asserted-by":"crossref","unstructured":"Charatan David Li Sizhe Lester Tagliasacchi Andrea andSitzmann Vincent. \u201cpixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 19457\u2013194673.","DOI":"10.1109\/CVPR52733.2024.01840"},{"key":"e_1_2_6_13_2","unstructured":"Chen Qihua Ma Yue Wang Hongfa et al. \u201cFollow-your-canvas: Higher-resolution video outpainting with extensive content generation\u201d.arXiv preprint arXiv:2409.01055(2024). 3."},{"key":"e_1_2_6_14_2","first-page":"370","volume-title":"European Conference on Computer Vision","author":"Chen Yuedong","year":"2024"},{"key":"e_1_2_6_15_2","unstructured":"Chen Shen Zhou Jiale andLi Lei. \u201cOptimizing 3D Gaussian Splatting for Sparse Viewpoint Scene Reconstruction\u201d.arXiv preprint arXiv:2409.03213(2024) 3."},{"key":"e_1_2_6_16_2","doi-asserted-by":"crossref","unstructured":"Duan Yuanxing Wei Fangyin Dai Qiyu et al. \u201c4d-rotor gaussian splatting: towards efficient novel view synthesis for dynamic scenes\u201d.ACM SIGGRAPH 2024 Conference Papers.2024 1\u2013113.","DOI":"10.1145\/3641519.3657463"},{"key":"e_1_2_6_17_2","first-page":"14304","volume-title":"2021 IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Du Yilun","year":"2021"},{"key":"e_1_2_6_18_2","unstructured":"Fan Zhiwen Cong Wenyan Wen Kairun et al. \u201cInstantsplat: Unbounded sparse-view pose-free gaussian splatting in 40 seconds\u201d.arXiv preprint arXiv:2403.20309(2024) 3."},{"key":"e_1_2_6_19_2","doi-asserted-by":"crossref","unstructured":"Fan Fanda Guo Chaoxu Gong Litong et al. \u201cHierarchical masked 3d diffusion model for video outpainting\u201d.ACM International Conference on Multimedia.2023 7890\u20137900. 3.","DOI":"10.1145\/3581783.3612478"},{"key":"e_1_2_6_20_2","first-page":"75468","article-title":"Cat3d: Create anything in 3d with multi-view diffusion models","volume":"37","author":"Gao Ruiqi","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_6_21_2","doi-asserted-by":"crossref","unstructured":"Guedon AntoineandLepetit Vincent. \u201cSugar: Surface-aligned gaussian splatting for efficient 3d mesh reconstruction and high-quality mesh rendering\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 5354\u20135363. 3.","DOI":"10.1109\/CVPR52733.2024.00512"},{"key":"e_1_2_6_22_2","unstructured":"Gu Bohai Luo Hao Guo Song andDong Peiran. \u201cAdvanced Video Inpainting Using Optical Flow-Guided Efficient Diffusion\u201d.arXiv preprint arXiv:2412.00857(2024) 3."},{"key":"e_1_2_6_23_2","first-page":"33768","article-title":"Monocular dynamic view synthesis: A reality check","volume":"35","author":"Gao Hang","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_6_24_2","doi-asserted-by":"crossref","unstructured":"Gu Zekai Yan Rui Lu Jiahao et al. \u201cDiffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control\u201d.arXiv preprint arXiv:2501.03847(2025) 2 6 7.","DOI":"10.1145\/3721238.3730607"},{"key":"e_1_2_6_25_2","unstructured":"Hong Wenyi Ding Ming Zheng Wendi et al. \u201cCogvideo: Large-scale pretraining for text-to-video generation via transformers\u201d.arXiv preprint arXiv:2205.15868(2022). 3."},{"key":"e_1_2_6_26_2","doi-asserted-by":"crossref","unstructured":"Hu Wenbo Gao Xiangjun Li Xiaoyu et al. \u201cDepthcrafter: Generating consistent long depth sequences for open-world videos\u201d.arXiv preprint arXiv:2409.02095(2024) 2.","DOI":"10.1109\/CVPR52734.2025.00193"},{"key":"e_1_2_6_27_2","doi-asserted-by":"crossref","unstructured":"Huang Ziqi He Yinan Yu Jiashuo et al. \u201cVBench: Comprehensive Benchmark Suite for Video Generative Models\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 21807\u2013218186 7.","DOI":"10.1109\/CVPR52733.2024.02060"},{"key":"e_1_2_6_28_2","unstructured":"Huang Jiaxin Miao Sheng Yang BangBnag et al. \u201cVivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting\u201d.arXiv preprint arXiv:2504.11092(2025) 2\u20134 7."},{"key":"e_1_2_6_29_2","unstructured":"Ho JonathanandSalimans Tim. \u201cClassifier-free diffusion guidance\u201d.arXiv preprint arXiv:2207.12598(2022) 6."},{"key":"e_1_2_6_30_2","doi-asserted-by":"crossref","unstructured":"Haque Ayaan Tancik Matthew Efros Alexei A et al. \u201cInstruct-nerf2nerf: Editing 3d scenes with instructions\u201d.IEEE\/CVF International Conference on Computer Vision.2023 19740\u201319750.","DOI":"10.1109\/ICCV51070.2023.01808"},{"key":"e_1_2_6_31_2","unstructured":"He Hao Xu Yinghao Guo Yuwei et al. \u201cCameraCtrl: Enabling Camera Control for Text-to-Video Generation\u201d.arXiv preprint arXiv:2404.02101(2024) 3."},{"key":"e_1_2_6_32_2","doi-asserted-by":"crossref","unstructured":"Huang Binbin Yu Zehao Chen Anpei et al. \u201c2d Gaussian splatting for geometrically accurate radiance fields\u201d.ACM SIGGRAPH Conference.2024 1\u2013113.","DOI":"10.1145\/3641519.3657428"},{"key":"e_1_2_6_33_2","unstructured":"He Hao Yang Ceyuan Lin Shanchuan et al. \u201cCameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models\u201d.arXiv preprint arXiv:2503.10592(2025) 3."},{"key":"e_1_2_6_34_2","doi-asserted-by":"crossref","unstructured":"Irshad Muhammad Zubair Kollar Thomas Laskey Michael et al. \u201cCenterSnap: Single-Shot Multi-Object 3D Shape Reconstruction and Categorical 6D Pose and Size Estimation\u201d.IEEE International Conference on Robotics and Automation.2022 10632\u2013106403.","DOI":"10.1109\/ICRA46639.2022.9811799"},{"key":"e_1_2_6_35_2","unstructured":"Jeong Hyeonho Lee Suhyeon andYe Jong Chul. \u201cReangle-A-Video: 4D Video Generation as Video-to-Video Translation\u201d.arXiv preprint arXiv:2503.09151(2025) 3."},{"key":"e_1_2_6_36_2","doi-asserted-by":"crossref","unstructured":"Ji Longbin Zhong Lei Wei Pengfei andLi Changjian. \u201cPoseTraj: Pose-Aware Trajectory Control in Video Diffusion\u201d.arXiv preprint arXiv:2503.16068(2025) 3 8.","DOI":"10.1109\/CVPR52734.2025.02121"},{"key":"e_1_2_6_37_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3592433","article-title":"3D Gaussian Splatting for Real-Time Radiance Field Rendering","volume":"42","author":"Kerbl Bernhard","year":"2023","journal-title":"ACM Transactions on Graphics"},{"key":"e_1_2_6_38_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3072959.3073599","article-title":"Tanks and Temples: Benchmarking Large-Scale Scene Reconstruction","volume":"36","author":"Knapitsch Arno","year":"2017","journal-title":"ACM Transactions on Graphics"},{"key":"e_1_2_6_39_2","unstructured":"Kingma Diederik P Welling Max et al.Auto-encoding variational bayes.20133."},{"key":"e_1_2_6_40_2","unstructured":"Kong Lingdong Yang Wesley Mei Jianbiao et al. \u201c3D and 4D world modeling: A survey\u201d.arXiv preprint arXiv:2509.07996(2025) 2."},{"key":"e_1_2_6_41_2","doi-asserted-by":"crossref","DOI":"10.1017\/9781108182874","volume-title":"Virtual reality","author":"LaValle Steven M.","year":"2023"},{"key":"e_1_2_6_42_2","doi-asserted-by":"crossref","unstructured":"Lee Minhyeok Cho Suhwan Shin Chajin et al. \u201cVideo diffusion models are strong video inpainter\u201d.AAAI Conference on Artificial Intelligence.2025 4526\u20134533. 3.","DOI":"10.1609\/aaai.v39i4.32477"},{"key":"e_1_2_6_43_2","unstructured":"Liu Tianqi Huang Zihao Chen Zhaoxi et al. \u201cFree4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency\u201d.arXiv preprint arXiv:2503.20785(2025) 3."},{"key":"e_1_2_6_44_2","unstructured":"Liang Ao Kong Lingdong Yan Tianyi et al. \u201cWorldLens: Full-spectrum evaluations of driving world models in real world\u201d.arXiv preprint arXiv:2512.10958(2025). 2."},{"key":"e_1_2_6_45_2","unstructured":"Li Yaowei Li Lingen Zhang Zhaoyang et al. \u201cBlobCtrl: A Unified and Flexible Framework for Element-level Image Generation and Editing\u201d.arXiv preprint arXiv:2503.13434(2025) 3."},{"key":"e_1_2_6_46_2","doi-asserted-by":"crossref","unstructured":"Li Zhengqi Niklaus Simon Snavely Noah andWang Oliver. \u201cNeural scene flow fields for space-time view synthesis of dynamic scenes\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2021 6498\u201365083.","DOI":"10.1109\/CVPR46437.2021.00643"},{"key":"e_1_2_6_47_2","unstructured":"Liu Kunhao Shao Ling andLu Shijian. \u201cNovel View Extrapolation with Video Diffusion Priors\u201d.arXiv preprint arXiv:2411.14208(2024) 3."},{"key":"e_1_2_6_48_2","unstructured":"Liu Fangfu Sun Wenqiang Wang Hanyang et al. \u201cReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model\u201d.arXiv preprint arXiv:2408.16767(2024) 3."},{"key":"e_1_2_6_49_2","doi-asserted-by":"crossref","unstructured":"Li Zhengqi Wang Qianqian Cole Forrester et al. \u201cDynibar: Neural dynamic image-based rendering\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2023 4273\u201342843.","DOI":"10.1109\/CVPR52729.2023.00416"},{"key":"e_1_2_6_50_2","doi-asserted-by":"crossref","unstructured":"Liu Ruoshi Wu Rundi Van Hoorick Basile et al. \u201cZero-1-to-3:Zero\u2013shotone image to 3d object\u201d.IEEE\/CVF International Conference on Computer Vision.2023 9298\u201393093.","DOI":"10.1109\/ICCV51070.2023.00853"},{"key":"e_1_2_6_51_2","unstructured":"Li Xiaowen Xue Haolan Ren Peiran andBo Liefeng. \u201cDiffuEraser: A Diffusion Model for Video Inpainting\u201d.arXiv preprint arXiv:2501.10018(2025) 3."},{"key":"e_1_2_6_52_2","doi-asserted-by":"crossref","unstructured":"Lu Tao Yu Mulin Xu Linning et al. \u201cScaffold-gs: Structured 3d gaussians for view-adaptive rendering\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 20654\u2013206643.","DOI":"10.1109\/CVPR52733.2024.01952"},{"key":"e_1_2_6_53_2","doi-asserted-by":"crossref","unstructured":"Li Jiahe Zhang Jiawei Bai Xiao et al. \u201cDngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normalization\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 20775\u2013207853.","DOI":"10.1109\/CVPR52733.2024.01963"},{"key":"e_1_2_6_54_2","unstructured":"Ma Baorui Gao Huachen Deng Haoge et al. \u201cYou See it You Got it: Learning 3D Creation on Pose-Free Videos at Scale\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.20252\u20137."},{"key":"e_1_2_6_55_2","unstructured":"Miao Sheng Huang Jiaxin Bai Dongfeng et al. \u201cEfficient Depth-Guided Urban View Synthesis\u201d.arXiv preprint arXiv:2407.12395(2024) 3."},{"key":"e_1_2_6_56_2","unstructured":"Meyer MaxwellandSpruyt Jack. \u201cBEN: Using Confidence-Guided Matting for Dichotomous Image Segmentation\u201d.arXiv preprint arXiv:2501.06230(2025) 5."},{"key":"e_1_2_6_57_2","doi-asserted-by":"crossref","unstructured":"Niemeyer Michael Barron Jonathan T Mildenhall BEN et al. \u201cRegnerf: Regularizing neural radiance fields for view synthesis from sparse inputs\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2022 5480\u201354903.","DOI":"10.1109\/CVPR52688.2022.00540"},{"key":"e_1_2_6_58_2","doi-asserted-by":"crossref","unstructured":"Pumarola Albert Corona Enric Pons-Moll Gerard andMoreno-Noguer Francesc. \u201cD-nerf: Neural radiance fields for dynamic scenes\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2021 10318\u2013103273.","DOI":"10.1109\/CVPR46437.2021.01018"},{"key":"e_1_2_6_59_2","unstructured":"Park Byeongjun Go Hyojun Nam Hyelin et al. \u201cSteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric Steering\u201d.arXiv preprint arXiv:2503.12024(2025) 3."},{"key":"e_1_2_6_60_2","unstructured":"Poole Ben Jain Ajay Barron Jonathan T andMildenhall Ben. \u201cDreamfusion: Text-to-3d using 2d diffusion\u201d.arXiv preprint arXiv:2209.14988(2022) 3."},{"key":"e_1_2_6_61_2","unstructured":"Park Jangho Kwon Taesung andYe Jong Chul. \u201cZero4D: Training-Free 4D Video Generation From Single Video Using Off-the-Shelf Video Diffusion Model\u201d.arXiv preprint arXiv:2503.22622(2025) 3."},{"key":"e_1_2_6_62_2","doi-asserted-by":"crossref","unstructured":"Park Keunhong Sinha Utkarsh Barron Jonathan T et al. \u201cNerfies: Deformable Neural Radiance Fields\u201d.IEEE\/CVF International Conference on Computer Vision.2021 5865\u201358743.","DOI":"10.1109\/ICCV48922.2021.00581"},{"key":"e_1_2_6_63_2","doi-asserted-by":"crossref","unstructured":"Park Keunhong Sinha Utkarsh Hedman Peter et al. \u201cHypernerf: A higher-dimensional representation for topologically varying neural radiance fields\u201d.arXiv preprint arXiv:2106.13228(2021) 3.","DOI":"10.1145\/3478513.3480487"},{"key":"e_1_2_6_64_2","doi-asserted-by":"crossref","unstructured":"Rombach Robin Blattmann Andreas Lorenz Dominik et al. \u201cHigh-resolution image synthesis with latent diffusion models\u201d.Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition.2022 10684\u2013106953.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_2_6_65_2","doi-asserted-by":"crossref","unstructured":"Ren Xuanchi Shen Tianchang Huang Jiahui et al. \u201cGen3c: 3d-informed world-consistent video generation with precise camera control\u201d.arXiv preprint arXiv:2503.03751(2025) 2\u20134.","DOI":"10.1109\/CVPR52734.2025.00574"},{"key":"e_1_2_6_66_2","first-page":"80220","article-title":"Genwarp: Single image to novel views with semantic-preserving generative warping","volume":"37","author":"Seo Junyoung","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_6_67_2","unstructured":"Sargent Kyle Li Zizhang Shah Tanmay et al. \u201cZeroshot 360-degree view synthesis from a single real image\u201d.arXiv preprint arXiv:2310.17994(2023) 3."},{"key":"e_1_2_6_68_2","unstructured":"Song Jiaming Meng Chenlin andErmon Stefano. \u201cDenoising diffusion implicit models\u201d.arXiv preprint arXiv:2010.02502(2020) 6."},{"key":"e_1_2_6_69_2","unstructured":"Szymanowicz Stanislaw Zhang Jason Y. Srinivasan Pratul et al. \u201cBolt3D: Generating 3D Scenes in Seconds\u201d.arXiv preprint arXiv:2503.14445(2025) 3."},{"key":"e_1_2_6_70_2","first-page":"313","volume-title":"European Conference on Computer Vision","author":"Van Hoorick Basile","year":"2024"},{"key":"e_1_2_6_71_2","doi-asserted-by":"crossref","unstructured":"Wang Guangcong Chen Zhaoxi Loy Chen Change andLiu Ziwei. \u201cSparsenerf: Distilling depth ranking for few-shot novel view synthesis\u201d.IEEE\/CVF International Conference on Computer Vision.2023 9065\u201390763.","DOI":"10.1109\/ICCV51070.2023.00832"},{"key":"e_1_2_6_72_2","unstructured":"Wang Chaoyang Eckart Ben Lucey Simon andGallo Orazio. \u201cNeural trajectory fields for dynamic novel view synthesis\u201d.arXiv preprint arXiv:2105.05994(2021). 3."},{"key":"e_1_2_6_73_2","doi-asserted-by":"crossref","unstructured":"Wu Rundi Gao Ruiqi Poole Ben et al. \u201cCat4d: Create anything in 4d with multi-view video diffusion models\u201d.arXiv preprint arXiv:2411.18613(2024) 2 3.","DOI":"10.1109\/CVPR52734.2025.02427"},{"key":"e_1_2_6_74_2","doi-asserted-by":"crossref","unstructured":"Wang Shuzhe Leroy Vincent Cabon Yohann et al. \u201cDUSt3R: Geometric 3D Vision Made Easy\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 20697\u2013207093.","DOI":"10.1109\/CVPR52733.2024.01956"},{"key":"e_1_2_6_75_2","doi-asserted-by":"crossref","unstructured":"Wu Rundi Mildenhall Ben Henzler Philipp et al. \u201cReconfusion: 3d reconstruction with diffusion priors\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 21551\u2013215613.","DOI":"10.1109\/CVPR52733.2024.02036"},{"key":"e_1_2_6_76_2","doi-asserted-by":"crossref","first-page":"455","DOI":"10.1007\/s12599-020-00658-9","article-title":"Virtual reality","volume":"62","author":"Wohlgenannt Isabell","year":"2020","journal-title":"Business & Information Systems Engineering"},{"key":"e_1_2_6_77_2","first-page":"153","volume-title":"European Conference on Computer Vision","author":"Wang Fu-Yun","year":"2024"},{"key":"e_1_2_6_78_2","unstructured":"Wang Zirui Wu Shangzhe Xie Weidi et al. \u201cNeRF\u2013: Neural radiance fields without known camera parameters\u201d.arXiv preprint arXiv:2102.07064(2021) 3."},{"key":"e_1_2_6_79_2","doi-asserted-by":"crossref","unstructured":"Wu Guanjun Yi Taoran Fang Jiemin et al. \u201c4d gaussian splatting for real-time dynamic scene rendering\u201d.Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition.2024 20310\u2013203203 5.","DOI":"10.1109\/CVPR52733.2024.01920"},{"key":"e_1_2_6_80_2","unstructured":"Wang Qianqian Ye Vickie Gao Hang et al. \u201cShape of motion: 4d reconstruction from a single video\u201d.arXiv preprint arXiv:2407.13764(2024) 3 6 7."},{"key":"e_1_2_6_81_2","doi-asserted-by":"crossref","unstructured":"Xian Wenqi Huang Jia-Bin Kopf Johannes andKim Changil. \u201cSpace-time neural irradiance fields for free-viewpoint video\u201d.Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition.2021 9421\u201394313.","DOI":"10.1109\/CVPR46437.2021.00930"},{"key":"e_1_2_6_82_2","first-page":"273","volume-title":"2023 IEEE International Symposium on Mixed and Augmented Reality (ISMAR)","author":"Xin Yingye","year":"2023"},{"key":"e_1_2_6_83_2","doi-asserted-by":"crossref","unstructured":"Yu Zehao Chen Anpei Huang Binbin et al. \u201cMip-splatting: Alias-free 3d gaussian splatting\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 19447\u2013194563.","DOI":"10.1109\/CVPR52733.2024.01839"},{"key":"e_1_2_6_84_2","doi-asserted-by":"crossref","unstructured":"Yang Ziyi Gao Xinyu Zhou Wen et al. \u201cDeformable 3d gaussians for high-fidelity monocular dynamic scene reconstruction\u201d.Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition.2024 20331\u2013203413.","DOI":"10.1109\/CVPR52733.2024.01922"},{"key":"e_1_2_6_85_2","unstructured":"Yu Mark Hu Wenbo Xing Jinbo andShan Ying. \u201cTrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models\u201d.arXiv preprint arXiv:2503.05638(2025) 2\u20134 6 7."},{"key":"e_1_2_6_86_2","unstructured":"Yang Honghui Huang Di Yin Wei et al. \u201cDepth any video with scalable synthetic data\u201d.arXiv preprint arXiv:2410.10815(2024) 2 5."},{"key":"e_1_2_6_87_2","unstructured":"Yoon Jae Shin Kim Kihwan Gallo Orazio et al. \u201cNovel view synthesis of dynamic scenes with globally coherent depths from a monocular camera\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2020 5336\u201353453."},{"key":"e_1_2_6_88_2","unstructured":"Yu Hanyang Long Xiaoxiao andTan Ping. \u201cLM-Gaussian: Boost Sparse-view 3D Gaussian Splatting with Large Model Priors\u201d.arXiv preprint arXiv:2409.03456(2024) 3."},{"key":"e_1_2_6_89_2","unstructured":"Yan Yunzhi Lin Haotong Zhou Chenxu et al. \u201cStreet Gaussians for Modeling Dynamic Urban Scenes\u201d.arXiv preprint arXiv:2401.01339(2024) 3."},{"key":"e_1_2_6_90_2","doi-asserted-by":"crossref","unstructured":"Yang Jiawei Pavone Marco andWang Yue. \u201cFreenerf: Improving few-shot neural rendering with free frequency regularization\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2023 8254\u201382633.","DOI":"10.1109\/CVPR52729.2023.00798"},{"key":"e_1_2_6_91_2","unstructured":"Yang Zhuoyi Teng Jiayan Zheng Wendi et al. \u201cCogvideox: Text-to-video diffusion models with an expert transformer\u201d.arXiv preprint arXiv:2408.06072(2024) 3."},{"key":"e_1_2_6_92_2","unstructured":"Yao Chun-Han Xie Yiming Voleti Vikram et al. \u201cSV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation\u201d.arXiv preprint arXiv:2503.16396(2025) 3."},{"key":"e_1_2_6_93_2","unstructured":"Yu Wangbo Xing Jinbo Yuan Li et al. \u201cViewcrafter: Taming video diffusion models for high-fidelity novel view synthesis\u201d.arXiv preprint arXiv:2409.02048(2024) 2 3 6 7."},{"key":"e_1_2_6_94_2","unstructured":"Yang Zeyu Yang Hongye Pan Zijie andZhang Li. \u201cReal-time photorealistic dynamic scene representation and rendering with 4d gaussian splatting\u201d.arXiv preprint arXiv:2310.10642(2023) 3."},{"key":"e_1_2_6_95_2","unstructured":"Yang Ling Zhu Kaixin Tian Juanxi et al. \u201cWideRange4D: Enabling High-Quality 4D Reconstruction with Wide-Range Movements and Scenes\u201d.arXiv preprint arXiv:2503.13435(2025) 3."},{"key":"e_1_2_6_96_2","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1109\/45.666641","article-title":"Virtual reality","volume":"17","author":"Zheng JM","year":"1998","journal-title":"Ieee Potentials"},{"key":"e_1_2_6_97_2","unstructured":"Zhao Slue Hu Wenbo Cun Xiaodong et al. \u201cStereocrafter: Diffusion-based generation of long and high-fidelity stereoscopic 3d from monocular videos\u201d.arXiv preprint arXiv:2409.07447(2024) 3."},{"key":"e_1_2_6_98_2","doi-asserted-by":"crossref","unstructured":"Zhang Richard Isola Phillip Efros Alexei A et al. \u201cThe unreasonable effectiveness of deep features as a perceptual metric\u201d.Proceedings of the IEEE conference on computer vision and pattern recognition.2018 586\u20135956.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_2_6_99_2","unstructured":"Zhao Yuyang Lin Chung-Ching Lin Kevin et al. \u201cGenXD: Generating any 3D and 4D scenes\u201d.International Conference on Learning Representations.20252 3."},{"key":"e_1_2_6_100_2","first-page":"335","volume-title":"European Conference on Computer Vision","author":"Zhang Jiawei","year":"2024"},{"key":"e_1_2_6_101_2","doi-asserted-by":"crossref","unstructured":"Zhou Hongyu Shao Jiahao Xu Lu et al. \u201cHugs: Holistic urban 3d scene understanding via gaussian splatting\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 21336\u2013213453.","DOI":"10.1109\/CVPR52733.2024.02016"},{"key":"e_1_2_6_102_2","doi-asserted-by":"crossref","unstructured":"Zhou ZhizhuoandTulsiani Shubham. \u201cSparsefusion: Distilling view-conditioned diffusion for 3d reconstruction\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2023 12588\u2013125973.","DOI":"10.1109\/CVPR52729.2023.01211"},{"key":"e_1_2_6_103_2","doi-asserted-by":"crossref","unstructured":"Zhang Zhixing Wu Bichen Wang Xiaoyan et al. \u201cAvid: Any-length video inpainting with diffusion model\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 7162\u201371723.","DOI":"10.1109\/CVPR52733.2024.00684"},{"key":"e_1_2_6_104_2","doi-asserted-by":"crossref","unstructured":"Zi Bojia Zhao Shihao Qi Xianbiao et al. \u201cCococo: Improving text-guided video inpainting for better consistency controllability and compatibility\u201d.AAAI Conference on Artificial Intelligence.2025 11067\u2013110763.","DOI":"10.1609\/aaai.v39i10.33203"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70345","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70345","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70345","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T08:51:44Z","timestamp":1777107104000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70345"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,25]]},"references-count":103,"alternative-id":["10.1111\/cgf.70345"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70345","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,25]]},"assertion":[{"value":"2026-04-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70345"}}