{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T13:46:47Z","timestamp":1774878407937,"version":"3.50.1"},"reference-count":60,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T00:00:00Z","timestamp":1774828800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T00:00:00Z","timestamp":1774828800000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62572230"],"award-info":[{"award-number":["62572230"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Creating ultra\u2010high\u2010resolution spatially varying bidirectional reflectance functions (SVBRDFs) is critical for photorealistic 3D content creation, to faithfully represent fine\u2010scale surface details required for close\u2010up rendering. However, achieving 4K generation faces two key challenges: (1) the need to synthesize multiple reflectance maps at full resolution, which multiplies the pixel budget and imposes prohibitive memory and computational cost, and (2) the requirement to maintain strong pixellevel alignment across maps at 4K, which is particularly difficult when adapting pretrained models designed for the RGB image domain. We introduce HiMat, a diffusion\u2010based framework tailored for efficient and diverse 4K SVBRDF generation. To address the first challenge, HiMat generates in a high\u2010compression latent space via a DC\u2010AE and employs a pretrained diffusion transformer with linear attention to improve per\u2010map efficiency. To address the second challenge, we propose CrossStitch, a lightweight convolutional module that enforces cross\u2010map consistency without incurring the cost of global attention. Our experiments show that HiMat achieves high\u2010fidelity 4K SVBRDF generation with superior efficiency, structural consistency, and diversity compared to prior methods. Beyond materials, our framework also generalizes to related applications such as intrinsic decomposition.<\/jats:p>","DOI":"10.1111\/cgf.70343","type":"journal-article","created":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T12:52:09Z","timestamp":1774875129000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["HiMat: DiT\u2010based Ultra\u2010High Resolution SVBRDF Generation"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6170-7339","authenticated-orcid":false,"given":"Zixiong","family":"Wang","sequence":"first","affiliation":[{"name":"College of Computer Science Nankai University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4800-832X","authenticated-orcid":false,"given":"Jian","family":"Yang","sequence":"additional","affiliation":[{"name":"College of Computer Science Nankai University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3674-295X","authenticated-orcid":false,"given":"Yiwei","family":"Hu","sequence":"additional","affiliation":[{"name":"Adobe Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3808-6092","authenticated-orcid":false,"given":"Milo\u0161","family":"Ha\u0161an","sequence":"additional","affiliation":[{"name":"Adobe Research"},{"name":"NVIDIA Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8943-8364","authenticated-orcid":false,"given":"Beibei","family":"Wang","sequence":"additional","affiliation":[{"name":"Nanjing University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,3,30]]},"reference":[{"key":"e_1_2_8_2_2","unstructured":"BrooksT. PeeblesB. HolmesC. DePueW. GuoY. JingL. SchnurrD. TaylorJ. LuhmanT. LuhmanE. NgC. WangR. RameshA.:Video generation models as world simulators. URL:https:\/\/openai.com\/research\/video-generation-models-as-world-simulators. 3 4"},{"key":"e_1_2_8_3_2","unstructured":"ComaniciG. BieberE. SchaekermannM. PasupatI. SachdevaN. DhillonI. BlisteinM. RamO. ZhangD. RosenE. et al.: Gemini 2.5: Pushing the frontier with advanced reasoning multimodality long context and next generation agentic capabilities.arXiv preprint arXiv:2507.06261(2025). 6"},{"key":"e_1_2_8_4_2","unstructured":"ChenJ. CaiH. ChenJ. XieE. YangS. TangH. LiM. HanS.: Deep compression autoencoder for efficient highresolution diffusion models. InThe Thirteenth International Conference on Learning Representations(2025). 2 4"},{"key":"e_1_2_8_5_2","first-page":"74","volume-title":"European Conference on Computer Vision","author":"Chen J.","year":"2024"},{"key":"e_1_2_8_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/357290.357293"},{"key":"e_1_2_8_7_2","doi-asserted-by":"crossref","unstructured":"ChenM. WangY. HuD. ZhuP. GuoJ. GuoY.: Dtdmat: A comprehensive svbrdf dataset with detailed text descriptions. InThe 19th ACM SIGGRAPH International Conference on Virtual-Reality Continuum and its Applications in Industry(2024) pp.1\u201315. 5","DOI":"10.1145\/3703619.3706053"},{"key":"e_1_2_8_8_2","unstructured":"ChenJ. ZouD. HeW. ChenJ. XieE. HanS. CaiH.: Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space. InIEEE International Conference on Computer Vision (ICCV)(2025). 11"},{"key":"e_1_2_8_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201378"},{"issue":"4","key":"e_1_2_8_10_2","article-title":"Flexible svbrdf capture with a multi-image deep network","volume":"38","author":"Deschaintre V.","year":"2019","journal-title":"Computer Graphics Forum (Proceedings of the Eurographics Symposium on Rendering)"},{"key":"e_1_2_8_11_2","unstructured":"DuR. ChangD. HospedalesT. SongY.-Z. MaZ.: Demofusion: Democratising high-resolution image generation with no $$$. InCVPR(2024). 2 3"},{"key":"e_1_2_8_12_2","first-page":"8780","article-title":"Diffusion models beat gans on image synthesis","volume":"34","author":"Dhariwal P.","year":"2021","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_8_13_2","unstructured":"EsserP. KulalS. BlattmannA. EntezariR. M\u00fcllerJ. SainiH. LeviY. LorenzD. SauerA. BoeselF. et al.: Scaling rectified flow transformers for high-resolution image synthesis. InForty-first international conference on machine learning(2024). 3 4"},{"key":"e_1_2_8_14_2","first-page":"50742","article-title":"Dreamsim: Learning new dimensions of human visual similarity using synthetic data","volume":"36","author":"Fu S.","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_8_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3528223.3530173"},{"key":"e_1_2_8_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417779"},{"key":"e_1_2_8_17_2","doi-asserted-by":"crossref","unstructured":"HuV. T. BaumannS. A. GuiM. GrebenkovaO. MaP. SchusterbauerJ. OmmerB.: Zigma: A dit-style zigzag mamba diffusion model. InECCV(2024). 3","DOI":"10.1007\/978-3-031-72664-4_9"},{"key":"e_1_2_8_18_2","doi-asserted-by":"crossref","unstructured":"HuY. GuerreroP. HasanM. RushmeierH. DeschaintreV.: Generating Procedural Materials from Text or Image Prompts. InACM SIGGRAPH 2023 Conference Proceedings(2023). 2","DOI":"10.1145\/3588432.3591520"},{"key":"e_1_2_8_19_2","volume-title":"Pacific Graphics Short Papers and Posters","author":"He Z.","year":"2023"},{"key":"e_1_2_8_20_2","unstructured":"HesselJ. HoltzmanA. ForbesM. BrasR. L. ChoiY.: Clipscore: A reference-free evaluation metric for image captioning.arXiv preprint arXiv:2104.08718(2021). 7 9"},{"key":"e_1_2_8_21_2","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.15199"},{"key":"e_1_2_8_22_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i4.32456"},{"key":"e_1_2_8_23_2","unstructured":"KingmaD. P. WellingM.: Auto-Encoding Variational Bayes. InThe 2nd International Conference on Learning Representations(2014). 3"},{"key":"e_1_2_8_24_2","unstructured":"LabsB. F.:Flux.https:\/\/github.com\/black-forest-labs\/flux 2024. 3"},{"key":"e_1_2_8_25_2","unstructured":"LipmanY. ChenR. T. Q. Ben-HamuH. NickelM. LeM.: Flow matching for generative modeling. InThe Eleventh International Conference on Learning Representations(2023). 4"},{"key":"e_1_2_8_26_2","doi-asserted-by":"crossref","unstructured":"LavinA. GrayS.: Fast algorithms for convolutional neural networks. InProceedings of the IEEE conference on computer vision and pattern recognition(2016) pp.4013\u20134021. 2 5","DOI":"10.1109\/CVPR.2016.435"},{"key":"e_1_2_8_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2021.3084827"},{"key":"e_1_2_8_28_2","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.14980"},{"key":"e_1_2_8_29_2","unstructured":"MaX. DeschaintreV. Ha\u0161anM. LuanF. ZhouK. WuH. HuY.: Materialpicker: Multi-modal material generation with diffusion transformers.ACM Trans. Graph. (July2025). 2 3 4 5 9"},{"key":"e_1_2_8_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3618358"},{"key":"e_1_2_8_31_2","unstructured":"PhungH. DaoQ. DaoT. PhanH. MetaxasD. TranA.: Dimsum: Diffusion mamba - a scalable and unified spatial-frequency method for image generation. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems(2024). 3"},{"key":"e_1_2_8_32_2","unstructured":"PeeblesW. XieS.: Scalable diffusion models with transformers. InProceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)(October2023) pp.4195\u20134205. 3 4"},{"key":"e_1_2_8_33_2","unstructured":"QinQ. ZhuoL. XinY. DuR. LiZ. FuB. LuY. LiX. LiuD. ZhuX. BeddowW. MillonE. Victor PerezW. W. QiaoY. ZhangB. LiuX. LiH. XuC. GaoP.:Luminaimage 2.0: A unified and efficient image generative framework 2025. 3 5"},{"key":"e_1_2_8_34_2","unstructured":"RombachR. BlattmannA. LorenzD. EsserP. OmmerB.: High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE\/CVF conference on computer vision and pattern recognition(2022) pp.10684\u201310695. 3 4 10"},{"key":"e_1_2_8_35_2","unstructured":"RameshA. DhariwalP. NicholA. ChuC. ChenM.:Hierarchical text-conditional image generation with clip latents 2022. URL:https:\/\/arxiv.org\/abs\/2204.06125 arXiv:2204.06125. 3"},{"key":"e_1_2_8_36_2","unstructured":"RogozhnikovA.: Einops: Clear and reliable tensor manipulations with einstein-like notation. InInternational Conference on Learning Representations(2022). 5"},{"key":"e_1_2_8_37_2","unstructured":"RobertsM. RamapuramJ. RanjanA. KumarA. BautistaM. A. PaczanN. WebbR. SusskindJ. M.: Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding. InProceedings of the IEEE\/CVF international conference on computer vision(2021) pp.10912\u201310922. 10"},{"key":"e_1_2_8_38_2","first-page":"25278","article-title":"Laion-5b: An open large-scale dataset for training next generation image-text models","volume":"35","author":"Schuhmann C.","year":"2022","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_8_39_2","doi-asserted-by":"crossref","unstructured":"SartorS. PeersP.: Matfusion: a generative diffusion model for svbrdf capture. InACM SIGGRAPH Asia Conference Proceedings(December2023). URL:https:\/\/doi.org\/10.1145\/3610548.3618194. 6","DOI":"10.1145\/3610548.3618194"},{"key":"e_1_2_8_40_2","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.12721"},{"key":"e_1_2_8_41_2","unstructured":"TeamG. RiviereM. PathakS. SessaP. G. HardinC. BhupatirajuS. HussenotL. MesnardT. ShahriariB. Ram\u00e9A. et al.: Gemma 2: Improving open language models at a practical size.arXiv preprint arXiv:2408.00118(2024). 5"},{"key":"e_1_2_8_42_2","unstructured":"VecchioG. DeschaintreV.: Matsynth: A modern pbr materials dataset. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2024). 2 5 6 9"},{"key":"e_1_2_8_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3688830"},{"key":"e_1_2_8_44_2","unstructured":"VaswaniA. ShazeerN. ParmarN. UszkoreitJ. JonesL. GomezA. N. KaiserL. PolosukhinI.: Attention is all you need. InProceedings of the 31st International Conference on Neural Information Processing Systems(Red Hook NY USA 2017) NIPS'17 Curran Associates Inc. p.6000\u20136010. 2 3 4"},{"key":"e_1_2_8_45_2","doi-asserted-by":"crossref","unstructured":"VecchioG. SortinoR. PalazzoS. SpampinatoC.: Matfuse: Controllable material generation with diffusion models. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(June2024) pp.4429\u20134438. 2 3 7 8","DOI":"10.1109\/CVPR52733.2024.00424"},{"key":"e_1_2_8_46_2","unstructured":"WalterB. MarschnerS. R. LiH. TorranceK. E.: Microfacet models for refraction through rough surfaces.Rendering techniques 2007(2007) 18th. 4"},{"key":"e_1_2_8_47_2","unstructured":"WuH. ZhangZ. ZhangW. ChenC. LiaoL. LiC. GaoY. WangA. ZhangE. SunW. YanQ. MinX. ZhaiG. LinW.: Q-align: teaching lmms for visual scoring via discrete text-defined levels. InProceedings of the 41st International Conference on Machine Learning(2024) ICML'24 JMLR.org. 7"},{"key":"e_1_2_8_48_2","unstructured":"XieE. ChenJ. ChenJ. CaiH. TangH. LinY. ZhangZ. LiM. ZhuL. LuY. HanS.: SANA: Efficient highresolution text-to-image synthesis with linear diffusion transformers. InThe Thirteenth International Conference on Learning Representations(2025). 2 3 5 6"},{"key":"e_1_2_8_49_2","unstructured":"XieE. ChenJ. ZhaoY. YuJ. ZhuL. LinY. ZhangZ. LiM. ChenJ. CaiH. LiuB. ZhouD. HanS.: Sana 1.5: Efficient scaling of training-time and inference-time compute in linear diffusion transformer. InInternational Conference on Machine Learning(January2025). 3 5"},{"key":"e_1_2_8_50_2","unstructured":"XueB. GuarneraC. ZhaoS. MontazeriZ.: Reflectancefusion: Diffusion-based text to svbrdf generation. InEurographics Symposium on Rendering(2024) Eurographics Association. 2 6 7 8"},{"key":"e_1_2_8_51_2","doi-asserted-by":"crossref","unstructured":"XinL. ZhangZ. WeiJ. GaoW. GaoD.: Dreampbr: Text-driven generation of high-resolution svbrdf with multi-modal guidance.IEEE International Conference on Multimedia & Expo(ICME)(2025). 2","DOI":"10.1109\/ICME59968.2025.11210202"},{"key":"e_1_2_8_52_2","doi-asserted-by":"crossref","unstructured":"YuF. GuJ. LiZ. HuJ. KongX. WangX. HeJ. QiaoY. DongC.: Scaling up to excellence: Practicing model scaling for photo-realistic image restoration in the wild. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2024) pp.25669\u201325680. 8 9","DOI":"10.1109\/CVPR52733.2024.02425"},{"key":"e_1_2_8_53_2","unstructured":"YuR. LiuS. TanZ. WangX.: Ultra-resolution adaptation with ease.International Conference on Machine Learning(2025). 3"},{"key":"e_1_2_8_54_2","unstructured":"YangZ. TengJ. ZhengW. DingM. HuangS. XuJ. YangY. HongW. ZhangX. FengG. YinD. Yuxuan Zhang WangW. ChengY. XuB. GuX. DongY. TangJ.: Cogvideox: Text-to-video diffusion models with an expert transformer. InThe Thirteenth International Conference on Learning Representations(2025). 3 4"},{"key":"e_1_2_8_55_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.visinf.2023.12.001"},{"key":"e_1_2_8_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3626235"},{"key":"e_1_2_8_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/3641519.3657445"},{"key":"e_1_2_8_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3550469.3555403"},{"key":"e_1_2_8_59_2","doi-asserted-by":"crossref","unstructured":"ZhouX. Ha\u0161anM. DeschaintreV. GuerreroP. Hold-GeoffroyY. SunkavalliK. KalantariN. K.: Photomat: A material generator learned from single flash photos. InSIGGRAPH 2023 Conference Papers(2023). 2","DOI":"10.1145\/3588432.3591535"},{"key":"e_1_2_8_60_2","doi-asserted-by":"crossref","unstructured":"ZhangJ. HuangQ. LiuJ. GuoX. HuangD.: Diffusion-4k: Ultra-high-resolution image synthesis with latent diffusion models. InIEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(2025). 3 7","DOI":"10.1109\/CVPR52734.2025.02185"},{"key":"e_1_2_8_61_2","doi-asserted-by":"crossref","unstructured":"ZhangR. IsolaP. EfrosA. A. ShechtmanE. WangO.: The unreasonable effectiveness of deep features as a perceptual metric. InCVPR(2018). 4","DOI":"10.1109\/CVPR.2018.00068"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70343","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70343","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70343","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T12:52:21Z","timestamp":1774875141000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70343"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,30]]},"references-count":60,"alternative-id":["10.1111\/cgf.70343"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70343","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,30]]},"assertion":[{"value":"2026-03-30","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70343"}}