{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T08:16:43Z","timestamp":1783066603985,"version":"3.54.6"},"reference-count":121,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T00:00:00Z","timestamp":1783036800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"NSFC\/RGC Collaborative Research Scheme","award":["CRS_HKUST605\/25"],"award-info":[{"award-number":["CRS_HKUST605\/25"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2026,7,3]]},"abstract":"<jats:p>\n                    Recent progress has shown that video diffusion models (VDMs) can be repurposed to solve various multimodal graphics tasks. However, existing approaches predominantly train separate models for each specific problem setting. This practice locks models into fixed input-output mappings, and typically ignores the joint correlations across modalities. In this paper, we present\n                    <jats:bold>UniVidX<\/jats:bold>\n                    , a unified multimodal framework designed to leverage VDM priors to enable versatile video generation. Our goal is to (i) master diverse pixel-aligned tasks by formulating them as conditional generation problems within multimodal space, (ii) adapt to modality-specific distributions without compromising the backbone's native priors, and (iii) ensure cross-modal consistency during synthesis. Concretely, we propose three key designs: 1) Stochastic Condition Masking (SCM): by randomly partitioning modalities into clean conditions and noisy targets during training, we enable the model to learn omni-directional conditional generation rather than fixed mappings. 2) Decoupled Gated LoRA (DGL): we attach per-modality LoRAs and activate them when a modality serves as a generation target, thereby preserving the VDM's strong priors. 3) Cross-Modal Self-Attention (CMSA): we explicitly share keys\/values across modalities while maintaining modality-specific queries, facilitating information exchange and inter-modal alignment. We validate our framework by instantiating it in two domains: 1)\n                    <jats:italic toggle=\"yes\">UniVid-Intrinsic<\/jats:italic>\n                    for RGB videos and their intrinsic maps (albedo, irradiance, normal), and 2)\n                    <jats:italic toggle=\"yes\">UniVid-Alpha<\/jats:italic>\n                    for blended RGB videos and their constituent RGBA layers. Experimental results demonstrate that both models achieve performance competitive with state-of-the-art methods across distinct tasks. Notably, they exhibit robust generalization capabilities in in-the-wild scenarios, even when trained on limited datasets of fewer than 1k videos.\n                  <\/jats:p>","DOI":"10.1145\/3811304","type":"journal-article","created":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:05:51Z","timestamp":1783062351000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors"],"prefix":"10.1145","volume":"45","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-4693-2326","authenticated-orcid":false,"given":"Houyuan","family":"Chen","sequence":"first","affiliation":[{"name":"MMLAB, HKUST, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4240-3073","authenticated-orcid":false,"given":"Hong","family":"Li","sequence":"additional","affiliation":[{"name":"Beihang University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-9865-4105","authenticated-orcid":false,"given":"Xianghao","family":"Kong","sequence":"additional","affiliation":[{"name":"MMLAB, HKUST, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-5146-019X","authenticated-orcid":false,"given":"Tianrui","family":"Zhu","sequence":"additional","affiliation":[{"name":"Nanjing University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7525-0790","authenticated-orcid":false,"given":"Shaocong","family":"Xu","sequence":"additional","affiliation":[{"name":"Beijing Academy of Artificial Intelligence, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-2548-0485","authenticated-orcid":false,"given":"Weiqing","family":"Xiao","sequence":"additional","affiliation":[{"name":"Nanjing University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-1516-4083","authenticated-orcid":false,"given":"Yuwei","family":"Guo","sequence":"additional","affiliation":[{"name":"MMLAB, The Chinese University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7123-0220","authenticated-orcid":false,"given":"Chongjie","family":"Ye","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3503-5791","authenticated-orcid":false,"given":"Lvmin","family":"Zhang","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7903-581X","authenticated-orcid":false,"given":"Hao","family":"Zhao","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1004-7753","authenticated-orcid":false,"given":"Anyi","family":"Rao","sequence":"additional","affiliation":[{"name":"MMLAB, HKUST, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,3]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Intrinsic dimensionality explains the effectiveness of language model fine-tuning. arXiv preprint arXiv:2012.13255","author":"Aghajanyan Armen","year":"2020","unstructured":"Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta. 2020. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. arXiv preprint arXiv:2012.13255 (2020)."},{"key":"e_1_2_1_2_1","volume-title":"Semantic soft segmentation. TOG","author":"Aksoy Ya\u011fiz","year":"2018","unstructured":"Ya\u011fiz Aksoy, Tae-Hyun Oh, Sylvain Paris, Marc Pollefeys, and Wojciech Matusik. 2018. Semantic soft segmentation. TOG (2018)."},{"key":"e_1_2_1_3_1","volume-title":"Tunc Ozan Aydin, and Marc Pollefeys","author":"Aksoy Yagiz","year":"2017","unstructured":"Yagiz Aksoy, Tunc Ozan Aydin, and Marc Pollefeys. 2017. Designing effective inter-pixel information flow for natural image matting. In CVPR."},{"key":"e_1_2_1_4_1","volume-title":"Segdiff: Image segmentation with diffusion probabilistic models. arXiv preprint arXiv:2112.00390","author":"Amit Tomer","year":"2021","unstructured":"Tomer Amit, Eliya Nachmani, Tal Shaharbany, and Lior Wolf. 2021. Segdiff: Image segmentation with diffusion probabilistic models. arXiv preprint arXiv:2112.00390 (2021)."},{"key":"e_1_2_1_5_1","volume-title":"Davison","author":"Bae Gwangbin","year":"2024","unstructured":"Gwangbin Bae and Andrew J. Davison. 2024. Rethinking Inductive Biases for Surface Normal Estimation. In CVPR."},{"key":"e_1_2_1_6_1","unstructured":"Shuai Bai et al. 2025. Qwen3-VL Technical Report. arXiv preprint arXiv:2511.21631 (2025)."},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Jonathan T Barron and Jitendra Malik. 2013. Intrinsic scene properties from a single rgb-d image. In CVPR.","DOI":"10.1109\/CVPR.2013.10"},{"key":"e_1_2_1_8_1","volume-title":"Intrinsic images in the wild. TOG","author":"Bell Sean","year":"2014","unstructured":"Sean Bell, Kavita Bala, and Noah Snavely. 2014. Intrinsic images in the wild. TOG (2014)."},{"key":"e_1_2_1_9_1","unstructured":"Yanrui Bin Wenbo Hu Haoyuan Wang Xinya Chen and Bing Wang. 2025. Normal-Crafter: Learning Temporally Consistent Normals from Video Diffusion Priors."},{"key":"e_1_2_1_10_1","unstructured":"Andreas Blattmann Tim Dockhorn Sumith Kulal Daniel Mendelevitch et al. 2023. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets. arXiv preprint arXiv:2311.15127 (2023)."},{"key":"e_1_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Nicolas Bonneel Balazs Kovacs Sylvain Paris and Kavita Bala. 2017. Intrinsic decompositions for image editing. In Computer graphics forum.","DOI":"10.1111\/cgf.13149"},{"key":"e_1_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Adrien Bousseau Sylvain Paris and Fr\u00e9do Durand. 2009. User-assisted intrinsic images. In SIGGRAPH Asia.","DOI":"10.1145\/1661412.1618476"},{"key":"e_1_2_1_13_1","unstructured":"Tim Brooks Bill Peebles Connor Homes Will DePue Yufei Guo Li Jing David Schnurr Joe Taylor et al. 2024. Video generation models as world simulators. OpenAI Technical Report (2024)."},{"key":"e_1_2_1_14_1","volume-title":"SIGGRAPH 2012 Course Notes.","author":"Burley Brent","year":"2012","unstructured":"Brent Burley and Walt Disney Animation Studios. 2012. Physically-based shading at disney. In SIGGRAPH 2012 Course Notes."},{"key":"e_1_2_1_15_1","unstructured":"Daniel J Butler Jonas Wulff Garrett B Stanley and Michael J Black. 2012. A naturalistic open source movie for optical flow evaluation. In ECCV."},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3630750","article-title":"Intrinsic image decomposition via ordinal shading","volume":"43","author":"Careaga Chris","year":"2023","unstructured":"Chris Careaga and Ya\u011fiz Aksoy. 2023. Intrinsic image decomposition via ordinal shading. ACM Transactions on Graphics 43, 1 (2023), 1\u201324.","journal-title":"ACM Transactions on Graphics"},{"key":"e_1_2_1_17_1","volume-title":"Colorful diffuse intrinsic image decomposition in the wild. TOG","author":"Careaga Chris","year":"2024","unstructured":"Chris Careaga and Ya\u011fiz Aksoy. 2024. Colorful diffuse intrinsic image decomposition in the wild. TOG (2024)."},{"key":"e_1_2_1_18_1","volume-title":"Tom-net: Learning transparent object matting from a single image. In CVPR.","author":"Chen Guanying","year":"2018","unstructured":"Guanying Chen, Kai Han, and Kwan-Yee K Wong. 2018. Tom-net: Learning transparent object matting from a single image. In CVPR."},{"key":"e_1_2_1_19_1","volume-title":"Real-time edge-aware image processing with the bilateral grid. TOG","author":"Chen Jiawen","year":"2007","unstructured":"Jiawen Chen, Sylvain Paris, and Fr\u00e9do Durand. 2007. Real-time edge-aware image processing with the bilateral grid. TOG (2007)."},{"key":"e_1_2_1_20_1","volume-title":"Video Depth Anything: Consistent Depth Estimation for Super-Long Videos. arXiv:2501.12375","author":"Chen Sili","year":"2025","unstructured":"Sili Chen, Hengkai Guo, Shengnan Zhu, Feihu Zhang, Zilong Huang, Jiashi Feng, and Bingyi Kang. 2025a. Video Depth Anything: Consistent Depth Estimation for Super-Long Videos. arXiv:2501.12375 (2025)."},{"key":"e_1_2_1_21_1","volume-title":"European Conference on Computer Vision. Springer, 450\u2013467","author":"Chen Xi","year":"2024","unstructured":"Xi Chen, Sida Peng, Dongchen Yang, Yuan Liu, Bowen Pan, Chengfei Lv, and Xiaowei Zhou. 2024. Intrinsicanything: Learning diffusion priors for inverse rendering under unknown illumination. In European Conference on Computer Vision. Springer, 450\u2013467."},{"key":"e_1_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Zhifei Chen Tianshuo Xu Wenhang Ge Leyi Wu Dongyu Yan Jing He Luozhou Wang Lu Zeng Shunsi Zhang and Ying-Cong Chen. 2025b. Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion. In CVPR.","DOI":"10.1109\/CVPR52734.2025.02468"},{"key":"e_1_2_1_23_1","unstructured":"Yusuf Dalva Yijun Li Qing Liu Nanxuan Zhao et al. 2024. LayerFusion: Harmonized Multi-Layer Text-to-Image Generation with Generative Priors. arXiv preprint arXiv:2412.04460 (2024)."},{"key":"e_1_2_1_24_1","volume-title":"PRISM: A Unified Framework for Photorealistic Reconstruction and Intrinsic Scene Modeling. arXiv preprint arXiv:2504.14219","author":"Dirik Alara","year":"2025","unstructured":"Alara Dirik, Tuanfeng Wang, Duygu Ceylan, Stefanos Zafeiriou, and Anna Fr\u00fchst\u00fcck. 2025. PRISM: A Unified Framework for Photorealistic Reconstruction and Intrinsic Scene Modeling. arXiv preprint arXiv:2504.14219 (2025)."},{"key":"e_1_2_1_25_1","volume-title":"Wan-Alpha: High-Quality Text-to-Video Generation with Alpha Channel. arXiv preprint arXiv:2509.24979","author":"Dong Haotian","year":"2025","unstructured":"Haotian Dong, Wenjing Wang, Chen Li, and Di Lin. 2025. Wan-Alpha: High-Quality Text-to-Video Generation with Alpha Channel. arXiv preprint arXiv:2509.24979 (2025)."},{"key":"e_1_2_1_26_1","volume-title":"Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans. In CVPR.","author":"Eftekhar Ainaz","year":"2021","unstructured":"Ainaz Eftekhar, Alexander Sax, Jitendra Malik, and Amir Zamir. 2021. Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans. In CVPR."},{"key":"e_1_2_1_27_1","doi-asserted-by":"crossref","unstructured":"Mikhail Erofeev Yury Gitman Dmitriy S Vatolin Alexey Fedorov and Jue Wang. 2015. Perceptually Motivated Benchmark for Video Matting.. In BMVC.","DOI":"10.5244\/C.29.99"},{"key":"e_1_2_1_28_1","unstructured":"Xiao Fu Wei Yin Mu Hu Kaixuan Wang Yuexin Ma Ping Tan Shaojie Shen Dahua Lin and Xiaoxiao Long. 2024. GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image. In ECCV."},{"key":"e_1_2_1_29_1","volume-title":"CAT3D: Create Anything in 3D with Multi-View Diffusion Models. NIPS","author":"Ruiqi","year":"2024","unstructured":"Ruiqi Gao*, Aleksander Holynski*, Philipp Henzler, Arthur Brussee, Ricardo Martin-Brualla, Pratul P. Srinivasan, Jonathan T. Barron, and Ben Poole*. 2024. CAT3D: Create Anything in 3D with Multi-View Diffusion Models. NIPS (2024)."},{"key":"e_1_2_1_30_1","doi-asserted-by":"crossref","unstructured":"Ioannis Gkioulekas Shuang Zhao Kavita Bala Todd Zickler and Anat Levin. 2013. Inverse volume rendering with material dictionaries. TOG (2013).","DOI":"10.1145\/2508363.2508377"},{"key":"e_1_2_1_31_1","unstructured":"Ming Gui Johannes Schusterbauer Ulrich Prestel Pingchuan Ma et al. 2024. DepthFM: Fast Monocular Depth Estimation with Flow Matching. arXiv preprint arXiv:2403.13788 (2024)."},{"key":"e_1_2_1_32_1","volume-title":"SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models. arXiv preprint arXiv:2311.16933","author":"Guo Yuwei","year":"2023","unstructured":"Yuwei Guo, Ceyuan Yang, Anyi Rao, Maneesh Agrawala, Dahua Lin, and Bo Dai. 2023. SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models. arXiv preprint arXiv:2311.16933 (2023)."},{"key":"e_1_2_1_33_1","doi-asserted-by":"crossref","unstructured":"Yuwei Guo Ceyuan Yang Anyi Rao Chenlin Meng Omer Bar-Tal Shuangrui Ding Maneesh Agrawala Dahua Lin and Bo Dai. 2025. Keyframe-Guided Creative Video Inpainting. In CVPR.","DOI":"10.1109\/CVPR52734.2025.01214"},{"key":"e_1_2_1_34_1","volume-title":"LumiX: Structured and Coherent Text-to-Intrinsic Generation. arXiv preprint arXiv:2512.02781","author":"Han Xu","year":"2025","unstructured":"Xu Han, Biao Zhang, Xiangjun Tang, Xianzhi Li, and Peter Wonka. 2025. LumiX: Structured and Coherent Text-to-Intrinsic Generation. arXiv preprint arXiv:2512.02781 (2025)."},{"key":"e_1_2_1_35_1","volume-title":"Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction. arXiv preprint arXiv:2409.18124","author":"He Jing","year":"2025","unstructured":"Jing He, Haodong Li, Wei Yin, Yixun Liang, et al. 2025. Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction. arXiv preprint arXiv:2409.18124 (2025)."},{"key":"e_1_2_1_36_1","doi-asserted-by":"crossref","unstructured":"Lukas H\u00f6llein Alja\u017e Bo\u017ei\u0107 Norman M\u00fcller David Novotny Hung-Yu Tseng Christian Richardt Michael Zollh\u00f6fer and Matthias Nie\u00dfner. 2024. Viewdiff: 3d-consistent image generation with text-to-image models. In CVPR.","DOI":"10.1109\/CVPR52733.2024.00482"},{"key":"e_1_2_1_37_1","volume-title":"CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers. arXiv preprint arXiv:2205.15868","author":"Hong Wenyi","year":"2022","unstructured":"Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. 2022. CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers. arXiv preprint arXiv:2205.15868 (2022)."},{"key":"e_1_2_1_38_1","unstructured":"Edward J Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In ICLR."},{"key":"e_1_2_1_39_1","unstructured":"Wenbo Hu Xiangjun Gao Xiaoyu Li Sijie Zhao Xiaodong Cun Yong Zhang Long Quan and Ying Shan. 2025. DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos. In CVPR."},{"key":"e_1_2_1_40_1","volume-title":"UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation. arXiv preprint arXiv:2512.07831","author":"Huang Jiehui","year":"2025","unstructured":"Jiehui Huang, Yuechen Zhang, Xu He, Yuan Gao, Zhi Cen, Bin Xia, Yan Zhou, Xin Tao, Pengfei Wan, and Jiaya Jia. 2025. UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation. arXiv preprint arXiv:2512.07831 (2025)."},{"key":"e_1_2_1_41_1","doi-asserted-by":"crossref","unstructured":"Wei-Lun Huang and Ming-Sui Lee. 2023. End-to-End Video Matting With Trimap Propagation. In CVPR.","DOI":"10.1109\/CVPR52729.2023.01378"},{"key":"e_1_2_1_42_1","volume-title":"Vbench: Comprehensive benchmark suite for video generative models. In CVPR.","author":"Huang Ziqi","year":"2024","unstructured":"Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. 2024. Vbench: Comprehensive benchmark suite for video generative models. In CVPR."},{"key":"e_1_2_1_43_1","volume-title":"Abhinav Shrivastava, and Joon-Young Lee.","author":"Huynh Chuong","year":"2024","unstructured":"Chuong Huynh, Seoung Wug Oh,, Abhinav Shrivastava, and Joon-Young Lee. 2024. MaGGIe: Masked Guided Gradual Human Instance Matting. In CVPR."},{"key":"e_1_2_1_44_1","volume-title":"Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi.","author":"Ilharco Gabriel","year":"2022","unstructured":"Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2022. Editing models with task arithmetic. arXiv preprint arXiv:2212.04089 (2022)."},{"key":"e_1_2_1_45_1","volume-title":"Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction. arXiv preprint arXiv:2504.07961","author":"Jiang Zeren","year":"2025","unstructured":"Zeren Jiang, Chuanxia Zheng, Iro Laina, Diane Larlus, and Andrea Vedaldi. 2025. Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction. arXiv preprint arXiv:2504.07961 (2025)."},{"key":"e_1_2_1_46_1","volume-title":"Rodrigo Caye Daudt, and Konrad Schindler","author":"Ke Bingxin","year":"2024","unstructured":"Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger, Rodrigo Caye Daudt, and Konrad Schindler. 2024. Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation. In CVPR."},{"key":"e_1_2_1_47_1","volume-title":"Lau","author":"Ke Zhanghan","year":"2022","unstructured":"Zhanghan Ke, Jiayu Sun, Kaican Li, Qiong Yan, and Rynson W.H. Lau. 2022. MODNet: Real-Time Trimap-Free Portrait Matting via Objective Decomposition. In AAAI."},{"key":"e_1_2_1_48_1","volume-title":"IntrinsiX: High-Quality PBR Generation using Image Priors. NIPS","author":"Kocsis Peter","year":"2025","unstructured":"Peter Kocsis, Lukas H\u00f6llein, and Matthias Nie\u00dfner. 2025. IntrinsiX: High-Quality PBR Generation using Image Priors. NIPS (2025)."},{"key":"e_1_2_1_49_1","volume-title":"Intrinsic Image Diffusion for Indoor Single-view Material Estimation. CVPR","author":"Kocsis Peter","year":"2024","unstructured":"Peter Kocsis, Vincent Sitzmann, and Matthias Nie\u00dfner. 2024. Intrinsic Image Diffusion for Indoor Single-view Material Estimation. CVPR (2024)."},{"key":"e_1_2_1_50_1","volume-title":"Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603","author":"Kong Weijie","year":"2024","unstructured":"Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. 2024. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603 (2024)."},{"key":"e_1_2_1_51_1","unstructured":"Duong H. Le Tuan Pham Sangho Lee Christopher Clark et al. 2024. One Diffusion to Generate Them All. arXiv preprint arXiv:2411.16318 (2024)."},{"key":"e_1_2_1_52_1","volume-title":"Generative Omnimatte: Learning to Decompose Video into Layers. In CVPR.","author":"Lee Yao-Chih","year":"2025","unstructured":"Yao-Chih Lee, Erika Lu, Sarah Rumbley, Michal Geyer, Jia-Bin Huang, Tali Dekel, and Forrester Cole. 2025. Generative Omnimatte: Learning to Decompose Video into Layers. In CVPR."},{"key":"e_1_2_1_53_1","volume-title":"Computer graphics forum","author":"Lettry Louis","unstructured":"Louis Lettry, Kenneth Vanhoey, and Luc Van Gool. 2018. Unsupervised deep singleimage intrinsic decomposition using illumination-varying image sequences. In Computer graphics forum, Vol. 37. Wiley Online Library, 409\u2013419."},{"key":"e_1_2_1_54_1","volume-title":"A closed-form solution to natural image matting. TPAMI","author":"Levin Anat","year":"2007","unstructured":"Anat Levin, Dani Lischinski, and Yair Weiss. 2007. A closed-form solution to natural image matting. TPAMI (2007)."},{"key":"e_1_2_1_55_1","volume-title":"Spectral matting. TPAMI","author":"Levin Anat","year":"2008","unstructured":"Anat Levin, Alex Rav-Acha, and Dani Lischinski. 2008. Spectral matting. TPAMI (2008)."},{"key":"e_1_2_1_56_1","volume-title":"Vmformer: End-to-end video matting with transformer. In WACV.","author":"Li Jiachen","year":"2024","unstructured":"Jiachen Li, Vidit Goel, Marianna Ohanyan, Shant Navasardyan, Yunchao Wei, and Humphrey Shi. 2024a. Vmformer: End-to-end video matting with transformer. In WACV."},{"key":"e_1_2_1_57_1","unstructured":"Jiachen Li Jitesh Jain and Humphrey Shi. 2024b. Matting anything. In CVPR."},{"key":"e_1_2_1_58_1","volume-title":"Tensosdf: Roughness-aware tensorial representation for robust geometry and material reconstruction. TOG","author":"Li Jia","year":"2024","unstructured":"Jia Li, Lu Wang, Lei Zhang, and Beibei Wang. 2024c. Tensosdf: Roughness-aware tensorial representation for robust geometry and material reconstruction. TOG (2024)."},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00255"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00942"},{"key":"e_1_2_1_61_1","doi-asserted-by":"crossref","unstructured":"Ruofan Liang Zan Gojcic Huan Ling Jacob Munkberg Jon Hasselgren Zhi-Hao Lin Jun Gao Alexander Keller Nandita Vijaykumar Sanja Fidler and Zian Wang. 2025. DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models.","DOI":"10.1109\/CVPR52734.2025.02428"},{"key":"e_1_2_1_62_1","unstructured":"Chung-Ching Lin Jiang Wang Kun Luo Kevin Lin Linjie Li Lijuan Wang and Zicheng Liu. 2023. Adaptive Human Matting for Dynamic Videos. In CVPR."},{"key":"e_1_2_1_63_1","unstructured":"Hongkai Lin Dingkang Liang Mingyang Du Xin Zhou and Xiang Bai. 2025. More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models. In NIPS."},{"key":"e_1_2_1_64_1","unstructured":"Shanchuan Lin Andrey Ryabtsev Soumyadip Sengupta Brian L Curless Steven M Seitz and Ira Kemelmacher-Shlizerman. 2021. Real-time high-resolution background matting. In CVPR."},{"key":"e_1_2_1_65_1","unstructured":"Shanchuan Lin Linjie Yang Imran Saleemi and Soumyadip Sengupta. 2022. Robust high-resolution video matting with temporal guidance. In WACV."},{"key":"e_1_2_1_66_1","volume-title":"Heli Ben-Hamu, Maximilian Nickel, and Matt Le.","author":"Lipman Yaron","year":"2022","unstructured":"Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2022. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747 (2022)."},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00331"},{"key":"e_1_2_1_68_1","doi-asserted-by":"crossref","unstructured":"Xiaoxiao Long Yuan-Chen Guo Cheng Lin Yuan Liu Zhiyang Dou Lingjie Liu Yuexin Ma Song-Hai Zhang Marc Habermann Christian Theobalt et al. 2023. Wonder3D: Single Image to 3D using Cross-Domain Diffusion. arXiv preprint arXiv:2310.15008 (2023).","DOI":"10.1109\/CVPR52733.2024.00951"},{"key":"e_1_2_1_69_1","volume-title":"Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983","author":"Loshchilov Ilya","year":"2016","unstructured":"Ilya Loshchilov and Frank Hutter. 2016. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983 (2016)."},{"key":"e_1_2_1_70_1","volume-title":"Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101","author":"Loshchilov Ilya","year":"2017","unstructured":"Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 (2017)."},{"key":"e_1_2_1_71_1","volume-title":"SIGGRAPH Conference Papers.","author":"Luo Jundan","unstructured":"Jundan Luo, Duygu Ceylan, Jae Shin Yoon, Nanxuan Zhao, Julien Philip, Anna Fr\u00fchst\u00fcck, Wenbin Li, Christian Richardt, and Tuanfeng Y. Wang. 2024. IntrinsicDiffusion: Joint Intrinsic Layers from Latent Diffusion Models. In SIGGRAPH Conference Papers."},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2020.3023565"},{"key":"e_1_2_1_73_1","volume-title":"Christian Schmidt, Daan de Geus, Alexander Hermans, and Bastian Leibe.","author":"Garcia Gonzalo Martin","year":"2025","unstructured":"Gonzalo Martin Garcia, Karim Abou Zeid, Christian Schmidt, Daan de Geus, Alexander Hermans, and Bastian Leibe. 2025. Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think. In WACV."},{"key":"e_1_2_1_74_1","unstructured":"Meituan LongCat Team Xunliang Cai Qilong Huang Zhuoliang Kang Hongyu Li et al. 2025. LongCat-Video Technical Report. arXiv preprint arXiv:2510.22200 (2025)."},{"key":"e_1_2_1_75_1","volume-title":"One4D: Unified 4D Generation and Reconstruction via Decoupled LoRA Control. arXiv preprint arXiv:2511.18922","author":"Mi Zhenxing","year":"2025","unstructured":"Zhenxing Mi, Yuxin Wang, and Dan Xu. 2025. One4D: Unified 4D Generation and Reconstruction via Decoupled LoRA Control. arXiv preprint arXiv:2511.18922 (2025)."},{"key":"e_1_2_1_76_1","volume-title":"T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. arXiv preprint arXiv:2302.08453","author":"Mou Chong","year":"2023","unstructured":"Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, Ying Shan, and Xiaohu Qie. 2023. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. arXiv preprint arXiv:2302.08453 (2023)."},{"key":"e_1_2_1_77_1","volume-title":"Training a Commercial-Level Video Generation Model in 200k. arXiv preprint arXiv:2503.09642","author":"Peng Xiangyu","year":"2025","unstructured":"Xiangyu Peng, Zangwei Zheng, Chenhui Shen, Tom Young, Xinying Guo, Binluo Wang, Hang Xu, Hongxin Liu, Mingyan Jiang, Wenjun Li, Yuhui Wang, Anbang Ye, Gang Ren, Qianran Ma, Wanying Liang, Xiang Lian, Xiwen Wu, Yuting Zhong, Zhuangyan Li, Chaoyu Gong, Guojun Lei, Leijun Cheng, Limin Zhang, Minghao Li, Ruijie Zhang, Silan Hu, Shijie Huang, Xiaokang Wang, Yuanheng Zhao, Yuqi Wang, Ziang Wei, and Yang You. 2025. Open-Sora 2.0: Training a Commercial-Level Video Generation Model in 200k. arXiv preprint arXiv:2503.09642 (2025)."},{"key":"e_1_2_1_78_1","volume-title":"Caiming Xiong, Silvio Savarese, et al.","author":"Qin Can","year":"2023","unstructured":"Can Qin, Shu Zhang, Ning Yu, Yihao Feng, Xinyi Yang, Yingbo Zhou, Huan Wang, Juan Carlos Niebles, Caiming Xiong, Silvio Savarese, et al. 2023. UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild. arXiv preprint arXiv:2305.11147 (2023)."},{"key":"e_1_2_1_79_1","doi-asserted-by":"crossref","unstructured":"Christoph Rhemann Carsten Rother Jue Wang Margrit Gelautz Pushmeet Kohli and Pamela Rott. 2009. A perceptually motivated online benchmark for image matting. In CVPR.","DOI":"10.1109\/CVPR.2009.5206503"},{"key":"e_1_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00869"},{"key":"e_1_2_1_81_1","doi-asserted-by":"crossref","unstructured":"Soumyadip Sengupta Vivek Jayaram Brian Curless Steven M Seitz and Ira Kemelmacher-Shlizerman. 2020. Background matting: The world is your green screen. In CVPR.","DOI":"10.1109\/CVPR42600.2020.00236"},{"key":"e_1_2_1_82_1","doi-asserted-by":"crossref","unstructured":"Xiaoyong Shen Aaron Hertzmann Jiaya Jia Sylvain Paris Brian Price Eli Shechtman and Ian Sachs. 2016. Automatic portrait segmentation for image stylization. In Computer Graphics Forum.","DOI":"10.1111\/cgf.12814"},{"key":"e_1_2_1_83_1","volume-title":"Dimitris Samaras, Nikos Paragios, and Iasonas Kokkinos.","author":"Shu Zhixin","year":"2018","unstructured":"Zhixin Shu, Mihir Sahasrabudhe, Riza Alp Guler, Dimitris Samaras, Nikos Paragios, and Iasonas Kokkinos. 2018. Deforming autoencoders: Unsupervised disentangling of shape and appearance. In ECCV."},{"key":"e_1_2_1_84_1","unstructured":"Zhixin Shu Ersin Yumer Sunil Hadap Kalyan Sunkavalli Eli Shechtman and Dimitris Samaras. 2017. Neural face editing with intrinsic image disentangling. In CVPR."},{"key":"e_1_2_1_85_1","unstructured":"Jianlin Su Yu Lu Shengfeng Pan Bo Wen and Yunfeng Liu. 2021. RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv:2104.09864"},{"key":"e_1_2_1_86_1","volume-title":"Ouroboros: Single-step Diffusion Models for Cycle-consistent Forward and Inverse Rendering. arXiv preprint arXiv:2508.14461","author":"Sun Shanlin","year":"2025","unstructured":"Shanlin Sun, Yifan Wang, Hanwen Zhang, Yifeng Xiong, Qin Ren, Ruogu Fang, Xiaohui Xie, and Chenyu You. 2025a. Ouroboros: Single-step Diffusion Models for Cycle-consistent Forward and Inverse Rendering. arXiv preprint arXiv:2508.14461 (2025)."},{"key":"e_1_2_1_87_1","volume-title":"Single image portrait relighting. TOG","author":"Sun Tiancheng","year":"2019","unstructured":"Tiancheng Sun, Jonathan T Barron, Yun-Ta Tsai, Zexiang Xu, Xueming Yu, Graham Fyffe, Christoph Rhemann, Jay Busch, Paul E Debevec, and Ravi Ramamoorthi. 2019. Single image portrait relighting. TOG (2019)."},{"key":"e_1_2_1_88_1","volume-title":"UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation. arXiv preprint arXiv:2505.24521","author":"Sun Yang-Tian","year":"2025","unstructured":"Yang-Tian Sun, Xin Yu, Zehuan Huang, Yi-Hua Huang, Yuan-Chen Guo, Ziyi Yang, Yan-Pei Cao, and Xiaojuan Qi. 2025b. UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation. arXiv preprint arXiv:2505.24521 (2025)."},{"key":"e_1_2_1_89_1","doi-asserted-by":"crossref","unstructured":"Jingwei Tang Yagiz Aksoy Cengiz Oztireli Markus Gross and Tunc Ozan Aydin. 2019. Learning-based sampling for natural image matting. In CVPR.","DOI":"10.1109\/CVPR.2019.00317"},{"key":"e_1_2_1_90_1","volume-title":"Cheng Perng Phoo, and Bharath Hariharan","author":"Tang Luming","year":"2023","unstructured":"Luming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo, and Bharath Hariharan. 2023. Emergent correspondence from image diffusion. NIPS (2023)."},{"key":"e_1_2_1_91_1","volume-title":"Unsupervised Zero-Shot Segmentation using Stable Diffusion. arXiv preprint arXiv:2308.12469","author":"Tian Junjiao","year":"2023","unstructured":"Junjiao Tian, Lavisha Aggarwal, Andrea Colaco, Zsolt Kira, and Mar Gonzalez-Franco. 2023. Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion. arXiv preprint arXiv:2308.12469 (2023)."},{"key":"e_1_2_1_92_1","volume-title":"Wan: Open and Advanced Large-Scale Video Generative Models. arXiv preprint arXiv:2503.20314","author":"Wan Team","year":"2025","unstructured":"Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, et al. 2025. Wan: Open and Advanced Large-Scale Video Generative Models. arXiv preprint arXiv:2503.20314 (2025)."},{"key":"e_1_2_1_93_1","volume-title":"Spongecake: A layered microflake surface appearance model. TOG","author":"Wang Beibei","year":"2022","unstructured":"Beibei Wang, Wenhua Jin, Milo\u0161 Ha\u0161an, and Ling-Qi Yan. 2022. Spongecake: A layered microflake surface appearance model. TOG (2022)."},{"key":"e_1_2_1_94_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCP56744.2023.10233761"},{"key":"e_1_2_1_95_1","volume-title":"OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding. arXiv preprint arXiv:2504.10825","author":"Xi Dianbing","year":"2025","unstructured":"Dianbing Xi, Jiepeng Wang, Yuanzhi Liang, Xi Qi, Yuchi Huo, Rui Wang, Chi Zhang, and Xuelong Li. 2025a. OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding. arXiv preprint arXiv:2504.10825 (2025)."},{"key":"e_1_2_1_96_1","unstructured":"Dianbing Xi Jiepeng Wang Yuanzhi Liang Xi Qiu et al. 2025b. CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion. arXiv preprint arXiv:2511.21129 (2025)."},{"key":"e_1_2_1_97_1","volume-title":"What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? arXiv preprint arXiv:2403.06090","author":"Xu Guangkai","year":"2024","unstructured":"Guangkai Xu, Yongtao Ge, Mingyu Liu, Chengxiang Fan, Kangyang Xie, Zhiyue Zhao, Hao Chen, and Chunhua Shen. 2024a. What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? arXiv preprint arXiv:2403.06090 (2024)."},{"key":"e_1_2_1_98_1","volume-title":"GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors. arXiv preprint arXiv:2504.01016","author":"Xu Tian-Xing","year":"2025","unstructured":"Tian-Xing Xu, Xiangjun Gao, Wenbo Hu, Xiaoyu Li, Song-Hai Zhang, and Ying Shan. 2025a. GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors. arXiv preprint arXiv:2504.01016 (2025)."},{"key":"e_1_2_1_99_1","volume-title":"Jodi: Unification of Visual Generation and Understanding via Joint Modeling. arXiv preprint arXiv:2505.19084","author":"Xu Yifeng","year":"2025","unstructured":"Yifeng Xu, Zhenliang He, Meina Kan, Shiguang Shan, and Xilin Chen. 2025b. Jodi: Unification of Visual Generation and Understanding via Joint Modeling. arXiv preprint arXiv:2505.19084 (2025)."},{"key":"e_1_2_1_100_1","volume-title":"CtrLoRA: An Extensible and Efficient Framework for Controllable Image Generation. arXiv preprint arXiv:2410.09400","author":"Xu Yifeng","year":"2024","unstructured":"Yifeng Xu, Zhenliang He, Shiguang Shan, and Xilin Chen. 2024b. CtrLoRA: An Extensible and Efficient Framework for Controllable Image Generation. arXiv preprint arXiv:2410.09400 (2024)."},{"key":"e_1_2_1_101_1","volume-title":"Depth Any Video with Scalable Synthetic Data. arXiv preprint arXiv:2410.10815","author":"Yang Honghui","year":"2024","unstructured":"Honghui Yang, Di Huang, Wei Yin, Chunhua Shen, Haifeng Liu, Xiaofei He, Binbin Lin, Wanli Ouyang, and Tong He. 2024a. Depth Any Video with Scalable Synthetic Data. arXiv preprint arXiv:2410.10815 (2024)."},{"key":"e_1_2_1_102_1","volume-title":"Daniil Pakhomov, Mengwei Ren, Jianming Zhang, Zhe Lin, Cihang Xie, and Yuyin Zhou.","author":"Yang Jinrui","year":"2025","unstructured":"Jinrui Yang, Qing Liu, Yijun Li, Soo Ye Kim, Daniil Pakhomov, Mengwei Ren, Jianming Zhang, Zhe Lin, Cihang Xie, and Yuyin Zhou. 2025a. Generative Image Layer Decomposition with Visual Effects. CVPR."},{"key":"e_1_2_1_103_1","doi-asserted-by":"crossref","unstructured":"Peiqing Yang Shangchen Zhou Jixin Zhao Qingyi Tao and Chen Change Loy. 2025c. MatAnyone: Stable Video Matting with Consistent Memory Propagation. In CVPR.","DOI":"10.1109\/CVPR52734.2025.00684"},{"key":"e_1_2_1_104_1","volume-title":"Cross-Domain Diffusion for High-Fidelity 3D Generation From a Single Image. TPAMI","author":"Yang Yuxiao","year":"2025","unstructured":"Yuxiao Yang, Xiaoxiao Long, Zhiyang Dou, Cheng Lin, Yuan Liu, Qingsong Yan, Yuexin Ma, Haoqian Wang, Zhiqiang Wu, and Wei Yin. 2025b. Wonder3D++: Cross-Domain Diffusion for High-Fidelity 3D Generation From a Single Image. TPAMI (2025)."},{"key":"e_1_2_1_105_1","unstructured":"Zhuoyi Yang Jiayan Teng Wendi Zheng Ming Ding Shiyu Huang Jiazheng Xu Yuanming Yang Wenyi Hong Xiaohan Zhang Guanyu Feng et al. 2024b. CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer. arXiv preprint arXiv:2408.06072 (2024)."},{"key":"e_1_2_1_106_1","volume-title":"ViTMatte: Boosting image matting with pre-trained plain vision transformers. Information Fusion","author":"Yao Jingfeng","year":"2024","unstructured":"Jingfeng Yao, Xinggang Wang, Shusheng Yang, and Baoyuan Wang. 2024a. ViTMatte: Boosting image matting with pre-trained plain vision transformers. Information Fusion (2024)."},{"key":"e_1_2_1_107_1","volume-title":"Matte anything: Interactive natural image matting with segment anything model. Image and Vision Computing","author":"Yao Jingfeng","year":"2024","unstructured":"Jingfeng Yao, Xinggang Wang, Lang Ye, and Wenyu Liu. 2024b. Matte anything: Interactive natural image matting with segment anything model. Image and Vision Computing (2024)."},{"key":"e_1_2_1_108_1","volume-title":"Stablenormal: Reducing diffusion variance for stable and sharp normal. TOG","author":"Ye Chongjie","year":"2024","unstructured":"Chongjie Ye, Lingteng Qiu, Xiaodong Gu, Qi Zuo, Yushuang Wu, Zilong Dong, Liefeng Bo, Yuliang Xiu, and Xiaoguang Han. 2024. Stablenormal: Reducing diffusion variance for stable and sharp normal. TOG (2024)."},{"key":"e_1_2_1_109_1","unstructured":"Hao Yu Jiabo Zhan Zile Wang Jinglin Wang et al. 2025. OmniAlpha: A Sequence-to-Sequence Framework for Unified Multi-Task RGBA Generation. arXiv preprint arXiv:2511.20211 (2025)."},{"key":"e_1_2_1_110_1","volume-title":"Taskonomy: Disentangling task transfer learning. In CVPR.","author":"Zamir Amir R","year":"2018","unstructured":"Amir R Zamir, Alexander Sax, William Shen, Leonidas J Guibas, Jitendra Malik, and Silvio Savarese. 2018. Taskonomy: Disentangling task transfer learning. In CVPR."},{"key":"e_1_2_1_111_1","volume-title":"SIGGRAPH Conference Papers.","author":"Zeng Zheng","year":"2024","unstructured":"Zheng Zeng, Valentin Deschaintre, Iliyan Georgiev, Yannick Hold-Geoffroy, et al. 2024. RGB\u2194X: Image decomposition and synthesis using material- and lighting-aware diffusion models. In SIGGRAPH Conference Papers."},{"key":"e_1_2_1_112_1","volume-title":"JointNet: Extending Text-to-Image Diffusion for Dense Distribution Modeling. ICLR","author":"Zhang Jingyang","year":"2024","unstructured":"Jingyang Zhang, Shiwei Li, Yuanxun Lu, Tian Fang, David McKinnon, Yanghai Tsin, Long Quan, and Yao Yao. 2024. JointNet: Extending Text-to-Image Diffusion for Dense Distribution Modeling. ICLR (2024)."},{"key":"e_1_2_1_113_1","volume-title":"Physg: Inverse rendering with spherical gaussians for physics-based material editing and relighting. In CVPR.","author":"Zhang Kai","year":"2021","unstructured":"Kai Zhang, Fujun Luan, Qianqian Wang, Kavita Bala, and Noah Snavely. 2021. Physg: Inverse rendering with spherical gaussians for physics-based material editing and relighting. In CVPR."},{"key":"e_1_2_1_114_1","volume-title":"Transparent Image Layer Diffusion using Latent Transparency. TOG","author":"Zhang Lvmin","year":"2024","unstructured":"Lvmin Zhang and Maneesh Agrawala. 2024. Transparent Image Layer Diffusion using Latent Transparency. TOG (2024)."},{"key":"e_1_2_1_115_1","doi-asserted-by":"crossref","unstructured":"Lvmin Zhang Anyi Rao and Maneesh Agrawala. 2023. Adding Conditional Control to Text-to-Image Diffusion Models.","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"e_1_2_1_116_1","volume-title":"Diception: A generalist diffusion model for visual perceptual tasks. arXiv preprint arXiv:2502.17157","author":"Zhao Canyu","year":"2025","unstructured":"Canyu Zhao, Mingyu Liu, Huanyi Zheng, Muzhi Zhu, Zhiyue Zhao, Hao Chen, Tong He, and Chunhua Shen. 2025. Diception: A generalist diffusion model for visual perceptual tasks. arXiv preprint arXiv:2502.17157 (2025)."},{"key":"e_1_2_1_117_1","volume-title":"Open-sora: Democratizing efficient video production for all. arXiv preprint arXiv:2412.20404","author":"Zheng Zangwei","year":"2024","unstructured":"Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. 2024. Open-sora: Democratizing efficient video production for all. arXiv preprint arXiv:2412.20404 (2024)."},{"key":"e_1_2_1_118_1","unstructured":"Shangchen Zhou Chongyi Li Kelvin C.K Chan and Chen Change Loy. 2023. ProPainter: Improving Propagation and Transformer for Video Inpainting. In ICCV."},{"key":"e_1_2_1_119_1","doi-asserted-by":"publisher","DOI":"10.1145\/3550469.3555407"},{"key":"e_1_2_1_120_1","doi-asserted-by":"crossref","unstructured":"Jingsen Zhu Fujun Luan Yuchi Huo Zihao Lin Zhihua Zhong Dianbing Xi Rui Wang Hujun Bao Jiaxiang Zheng and Rui Tang. 2022b. Learning-based inverse rendering of complex indoor scenes with differentiable monte carlo raytracing. In Siggraph asia 2022 conference papers. 1\u20138.","DOI":"10.1145\/3550469.3555407"},{"key":"e_1_2_1_121_1","doi-asserted-by":"crossref","unstructured":"Junhao Zhuang Yanhong Zeng Wenran Liu Chun Yuan and Kai Chen. 2024. A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting. In ECCV.","DOI":"10.1007\/978-3-031-73636-0_12"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","deposited":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:25:38Z","timestamp":1783063538000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811304"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,3]]},"references-count":121,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,7,3]]}},"alternative-id":["10.1145\/3811304"],"URL":"https:\/\/doi.org\/10.1145\/3811304","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,3]]},"assertion":[{"value":"2026-01-15","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}