{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T11:32:21Z","timestamp":1776166341332,"version":"3.50.1"},"reference-count":53,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T00:00:00Z","timestamp":1776124800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T00:00:00Z","timestamp":1776124800000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100014188","name":"Ministry of Science and ICT, South Korea","doi-asserted-by":"publisher","award":["RS\u20102026\u201025486262"],"award-info":[{"award-number":["RS\u20102026\u201025486262"]}],"id":[{"id":"10.13039\/501100014188","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Reconstructing detailed geometry and realistic appearance from a single RGB image is essential yet fundamentally challenging due to inherent ambiguities such as occlusion, lighting variations, and texture\u2010geometry entanglement. While recent diffusion\u2010based generative models have significantly improved novel view synthesis, existing approaches suffer from two critical limitations: lack of cross\u2010view geometric consistency and insufficient cross\u2010domain semantic alignment. To address these issues, we introduce\n                    <jats:italic>\n                      U\n                      <jats:sc>ni<\/jats:sc>\n                      C\n                      <jats:sc>ross<\/jats:sc>\n                      3D\n                    <\/jats:italic>\n                    , a unified cross\u2010view and cross\u2010domain diffusion framework designed explicitly for consistent and physically coherent 3D generation.\n                    <jats:italic>\n                      U\n                      <jats:sc>ni<\/jats:sc>\n                      C\n                      <jats:sc>ross<\/jats:sc>\n                      3D\n                    <\/jats:italic>\n                    features two novel contributions: (1) a cross\u2010view latent regularization that enforces cross\u2010view geometric consistency across synthesized viewpoints by penalizing latent variance, and (2) a cross\u2010domain mutual information objective grounded in the physics of image formation, explicitly aligning synthesized color and normal maps. Extensive experiments demonstrate that\n                    <jats:italic>\n                      U\n                      <jats:sc>ni<\/jats:sc>\n                      C\n                      <jats:sc>ross<\/jats:sc>\n                      3D\n                    <\/jats:italic>\n                    achieves significantly improved view consistency and semantic alignment over state\u2010of\u2010the\u2010art methods and yields higher\u2010fidelity reconstructions, particularly under challenging textures and ambiguous viewpoints.\n                  <\/jats:p>","DOI":"10.1111\/cgf.70378","type":"journal-article","created":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T10:16:53Z","timestamp":1776161813000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["UniCross3D: Unified Cross\u2010View and Cross\u2010Domain Diffusion for Consistent Single\u2010Image 3D Generation"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-8742-3398","authenticated-orcid":false,"given":"U\u2010Chae","family":"Jun","sequence":"first","affiliation":[{"name":"Sookmyung Women's University  South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-9914-481X","authenticated-orcid":false,"given":"Jaeeun","family":"Ko","sequence":"additional","affiliation":[{"name":"Sookmyung Women's University  South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7622-0817","authenticated-orcid":false,"given":"Jiwoo","family":"Kang","sequence":"additional","affiliation":[{"name":"Sookmyung Women's University  South Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,4,14]]},"reference":[{"key":"e_1_2_7_2_2","unstructured":"Belghazi Mohamed Ishmael Baratin Aristide Rajeshwar Sai et al. \u201cMutual Information Neural Estimation\u201d.Proceedings of the International Conference on Machine Learning.20186."},{"key":"e_1_2_7_3_2","doi-asserted-by":"crossref","unstructured":"Boss Mark Huang Zixuan Vasishta Aaryaman andJampani Varun. \u201cSF3D: Stable fast 3D mesh reconstruction with uv-unwrapping and illumination disentanglement\u201d.Proceedings of the Computer Vision and Pattern Recognition Conference.2025 16240\u2013162504.","DOI":"10.1109\/CVPR52734.2025.01514"},{"issue":"8","key":"e_1_2_7_4_2","doi-asserted-by":"crossref","first-page":"1670","DOI":"10.1109\/TPAMI.2014.2377712","article-title":"Shape, Illumination, and Reflectance from Shading","volume":"37","author":"Barron Jonathan T","year":"2015","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_2_7_5_2","unstructured":"Chan Eric R. Lin Connor Z. Chan Matthew A. et al. \u201cEfficient Geometry-aware 3D Generative Adversarial Networks\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.20222."},{"key":"e_1_2_7_6_2","doi-asserted-by":"crossref","unstructured":"Caron Mathilde Touvron Hugo Misra Ishan et al. \u201cEmerging Properties in Self-Supervised Vision Transformers\u201d.Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV).2021 9650\u201396609.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"e_1_2_7_7_2","unstructured":"Choy Christopher Xu Danfei Gwak JunYoung et al. \u201c3D-R2N2: A Unified Approach for Single and Multi-view 3D Object Reconstruction\u201d.Proceedings of the European Conference on Computer Vision.20161."},{"key":"e_1_2_7_8_2","doi-asserted-by":"crossref","first-page":"1237","DOI":"10.1609\/aaai.v38i2.27886","article-title":"IT3D: Improved text-to-3D generation with explicit view synthesis","volume":"38","author":"Chen Yiwen","year":"2024","journal-title":"AAAI Conference on Artificial Intelligence."},{"key":"e_1_2_7_9_2","first-page":"11","volume":"9","author":"Downs Laura","year":"2022","journal-title":"International Conference on Robotics and Automation."},{"key":"e_1_2_7_10_2","doi-asserted-by":"crossref","unstructured":"Deitke Matt Schwenk Dustin Salvador Jordi et al. \u201cObjaverse: A Universe of Annotated 3D Objects\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2023 13142\u2013131537 12 15.","DOI":"10.1109\/CVPR52729.2023.01263"},{"key":"e_1_2_7_11_2","unstructured":"Fort Stanislav Hu Huiyi andLakshminarayanan Balaji. \u201cDeep Ensembles: A Loss Landscape Perspective\u201d.arXiv preprint arXiv:1912.02757(2019) 5."},{"key":"e_1_2_7_12_2","doi-asserted-by":"crossref","unstructured":"Feng Xiang Yu Chang Bi Zoubin et al. \u201cARM: Appearance reconstruction model for relightable 3D generation\u201d.Proceedings of the Computer Vision and Pattern Recognition Conference.2025 21425\u2013214374.","DOI":"10.1109\/CVPR52734.2025.01996"},{"key":"e_1_2_7_13_2","first-page":"1050","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Gal Yarin","year":"2016"},{"key":"e_1_2_7_14_2","doi-asserted-by":"crossref","unstructured":"Ge Wenhang Hu Tao Zhao Haoyu et al. \u201cRef-NeuS: Ambiguity-reduced neural implicit surface learning for multi-view reconstruction with reflection\u201d.IEEE\/CVF International Conference on Computer Vision.2023 4251\u2013426013.","DOI":"10.1109\/ICCV51070.2023.00392"},{"key":"e_1_2_7_15_2","unstructured":"Hjelm R. Devon Fedorov Alex Lavoie-Marchildon Samuel et al. \u201cLearning Deep Representations by Mutual Information Estimation and Maximization\u201d.Proceedings of the International Conference on Learning Representations.20196."},{"key":"e_1_2_7_16_2","unstructured":"Ho Jonathan Jain Ajay andAbbeel Pieter. \u201cDenoising Diffusion Probabilistic Models\u201d.Proceedings of the Advances in Neural Information Processing Systems.20201."},{"key":"e_1_2_7_17_2","doi-asserted-by":"crossref","DOI":"10.1016\/j.imavis.2024.105043","article-title":"EMOVA: Emotion-driven neural volumetric avatar","volume":"146","author":"Hwang Juheon","year":"2024","journal-title":"Image and Vision Computing"},{"key":"e_1_2_7_18_2","unstructured":"Heusel Martin Ramsauer Hubert Unterthiner Thomas et al. \u201cGANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium\u201d.Proceedings of the Advances in Neural Information Processing Systems.2017 6626\u201366377."},{"key":"e_1_2_7_19_2","unstructured":"Ho JonathanandSalimans Tim. \u201cClassifier-free diffusion guidance\u201d.arXiv preprint arXiv:2207.12598(2022) 9."},{"key":"e_1_2_7_20_2","unstructured":"Hong Yicong Zhang Kai Gu Jiuxiang et al. \u201cLRM: Large Reconstruction Model for Single Image to 3D\u201d.Proceedings of the International Conference on Learning Representations.20241\u20133 7 9."},{"key":"e_1_2_7_21_2","first-page":"439","volume-title":"European Conference on Computer Vision","author":"Jiang Chenhan","year":"2024"},{"issue":"4","key":"e_1_2_7_22_2","first-page":"139:1","article-title":"3D Gaussian Splatting for Real-Time Radiance Field Rendering","volume":"42","author":"Kerbl Bernhard","year":"2023","journal-title":"ACM Transactions on Graphics"},{"issue":"5","key":"e_1_2_7_23_2","doi-asserted-by":"crossref","first-page":"2858","DOI":"10.1109\/TSMC.2021.3054677","article-title":"Competitive learning of facial fitting and synthesis using UV energy","volume":"52","author":"Kang Jiwoo","year":"2021","journal-title":"IEEE Transactions on Systems, Man, and Cybernetics: Systems"},{"key":"e_1_2_7_24_2","unstructured":"Long Xiaoxiao Guo Yuan-Chen Lin Cheng et al. \u201cWonder3D: Single Image to 3D using Cross-Domain Diffusion\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.20241\u20133 7 9 11 13."},{"key":"e_1_2_7_25_2","unstructured":"Lin Chen-Hsuan Gao Jun Tang Luming et al. \u201cMagic3D: High-resolution Text-to-3D Content Creation\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.20231 2."},{"key":"e_1_2_7_26_2","unstructured":"Li Peng Liu Yuan Long Xiaoxiao et al. \u201cEra3D: HighResolution Multi-View Diffusion Using Efficient Row-Wise Attention\u201d.Advances in Neural Information Processing Systems.20241\u20133 7 9 11 13 14."},{"key":"e_1_2_7_27_2","unstructured":"Liu Yuan Lin Cheng Zeng Zijiao et al. \u201cSyncDreamer: Generating Multiview-consistent Images from a Single-view Image\u201d.Proceedings of the International Conference on Learning Representations.20241\u20133 7 9 11."},{"key":"e_1_2_7_28_2","unstructured":"Li Jiahao Tan Hao Zhang Kai et al. \u201cInstant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction Model\u201d.Proceedings of the International Conference on Learning Representations.20241\u20133."},{"key":"e_1_2_7_29_2","unstructured":"Liu Ruoshi Wu Rundi Van Hoorick Basile et al. \u201cZero-1-to-3: Zero-shot One Image to 3D Object\u201d.Proceedings of the IEEE\/CVF International Conference on Computer Vision.20231 3 7."},{"key":"e_1_2_7_30_2","first-page":"22226","article-title":"One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape Optimization","volume":"36","author":"Liu Minghua","year":"2023","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_2_7_31_2","doi-asserted-by":"crossref","unstructured":"Lin Jiantao Yang Xin Chen Meixi et al. \u201cKiss3DGen: Repurposing image diffusion models for 3D asset generation\u201d.Proceedings of the Computer Vision and Pattern Recognition Conference.2025 5870\u201358803 9.","DOI":"10.1109\/CVPR52734.2025.00551"},{"key":"e_1_2_7_32_2","first-page":"1","volume-title":"European Conference on Computer Vision","author":"Ma Zhiyuan","year":"2024"},{"key":"e_1_2_7_33_2","unstructured":"Podell Dustin English Zion Lacey Kyle et al. \u201cSDXL: Improving latent diffusion models for high-resolution image synthesis\u201d.arXiv preprint arXiv:2307.01952(2023) 7 9 10."},{"key":"e_1_2_7_34_2","unstructured":"Poole Ben Jain Ajay Barron Jonathan T. andMildenhall Ben. \u201cDreamFusion: Text-to-3D using 2D Diffusion\u201d.Proceedings of the International Conference on Learning Representations.20231 2 10."},{"key":"e_1_2_7_35_2","unstructured":"Poole Ben Ozair Sherjil van denOord Aaron et al. \u201cOn Variational Bounds of Mutual Information\u201d.Proceedings of the International Conference on Machine Learning.20196."},{"key":"e_1_2_7_36_2","doi-asserted-by":"crossref","unstructured":"Peebles WilliamandXie Saining. \u201cScalable diffusion models with transformers\u201d.Proceedings of the IEEE\/CVF international conference on computer vision.2023 4195\u201342053.","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"e_1_2_7_37_2","unstructured":"Qian Guocheng Mai Jinjie Hamdi Abdullah et al. \u201cMagic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors\u201d.Proceedings of the International Conference on Learning Representations.20241\u20133."},{"key":"e_1_2_7_38_2","unstructured":"Rombach Robin Blattmann Andreas Lorenz Dominik et al. \u201cHigh-Resolution Image Synthesis with Latent Diffusion Models\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.20221 9."},{"key":"e_1_2_7_39_2","unstructured":"Radford Alec Kim Jong Wook Hallacy Chris et al. \u201cLearning Transferable Visual Models from Natural Language Supervision\u201d.Proceedings of the International Conference on Machine Learning.2021 8748\u201387637."},{"key":"e_1_2_7_40_2","first-page":"36479","article-title":"Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding","volume":"35","author":"Saharia Chitwan","year":"2022","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_2_7_41_2","first-page":"6087","article-title":"Deep marching tetrahedra: A hybrid representation for high-resolution 3D shape synthesis","volume":"34","author":"Shen Tianchang","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_7_42_2","unstructured":"Song Jiaming Meng Chenlin andErmon Stefano. \u201cDenoising Diffusion Implicit Models\u201d.International Conference on Learning Representations.20219."},{"key":"e_1_2_7_43_2","unstructured":"Tang Jiaxiang Chen Zhaoxi Chen Xiaokang et al. \u201cLGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation\u201d.Proceedings of the European Conference on Computer Vision.20242 3."},{"key":"e_1_2_7_44_2","unstructured":"Van den Oord Aaron Li Yazhe andVinyals Oriol. \u201cRepresentation Learning with Contrastive Predictive Coding\u201d.arXiv preprint arXiv:1807.03748(2018) 6."},{"issue":"4","key":"e_1_2_7_45_2","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","article-title":"Image Quality Assessment: From Error Visibility to Structural Similarity","volume":"13","author":"Wang Zhou","year":"2004","journal-title":"IEEE Transactions on Image Processing"},{"key":"e_1_2_7_46_2","unstructured":"Wang Haochen Du Xiaodan Li Jiahao et al. \u201cScore Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D Generation\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.20231 3."},{"key":"e_1_2_7_47_2","unstructured":"Wu Kailu Liu Fangfu Cai Zhihan et al. \u201cUnique3D: High-quality and Efficient 3D Mesh Generation from a Single Image\u201d.Proceedings of the Advances in Neural Information Processing Systems.20243."},{"key":"e_1_2_7_48_2","unstructured":"Wang Peng Liu Lingjie Liu Yuan et al. \u201cNeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction\u201d.Proceedings of the Advances in Neural Information Processing Systems.20212 4 6."},{"key":"e_1_2_7_49_2","first-page":"3157","article-title":"UniSDF: Unifying neural representations for high-fidelity 3D reconstruction of complex scenes with reflections","volume":"37","author":"Wang Fangjinhua","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_7_50_2","doi-asserted-by":"crossref","unstructured":"Wang Zhengyi Wang Yikai Chen Yifei et al. \u201cCRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model\u201d.European Conference on Computer Vision.2024 57\u2013747 9 10.","DOI":"10.1007\/978-3-031-72751-1_4"},{"key":"e_1_2_7_51_2","unstructured":"Xu Jiale Cheng Weihao Gao Yiming et al. \u201cInstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models\u201d.arXiv preprint arXiv:2404.07191(2024) 2."},{"key":"e_1_2_7_52_2","unstructured":"Yang Haibo Chen Yang Pan Yingwei et al. \u201cHi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models\u201d.Proceedings of the ACM International Conference on Multimedia.20242."},{"key":"e_1_2_7_53_2","unstructured":"Zhang Richard Isola Phillip Efros Alexei A. et al. \u201cThe Unreasonable Effectiveness of Deep Features as a Perceptual Metric\u201d.IEEE\/CVF Conference on Computer Vision and Pattern Recognition.20187."},{"issue":"8","key":"e_1_2_7_54_2","doi-asserted-by":"crossref","first-page":"690","DOI":"10.1109\/34.784284","article-title":"Shape-from-Shading: A Survey","volume":"21","author":"Zhang Ruo","year":"1999","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70378","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70378","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70378","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T10:17:14Z","timestamp":1776161834000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70378"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,14]]},"references-count":53,"alternative-id":["10.1111\/cgf.70378"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70378","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,14]]},"assertion":[{"value":"2026-04-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70378"}}