{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T08:17:46Z","timestamp":1783066666957,"version":"3.54.6"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T00:00:00Z","timestamp":1783036800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"funder":[{"name":"Start-up fund from Tsinghua University"},{"name":"HKSAR Government under the ITSP-Platform grants","award":["ITS\/335\/23FP"],"award-info":[{"award-number":["ITS\/335\/23FP"]}]},{"name":"HKSAR Government under the ITSP-Platform grants","award":["ITS\/469\/24FP"],"award-info":[{"award-number":["ITS\/469\/24FP"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2026,7,3]]},"abstract":"<jats:p>\n                    In this paper, we study an under-explored but important factor of diffusion generative models, i.e., the combinatorial complexity. Data samples are generally high-dimensional, and for various structured generation tasks, additional attributes are combined to associate with data samples. We show that the space spanned by the combination of dimensions and attributes can be insufficiently covered by existing training schemes of diffusion generative models, potentially limiting test time performance. We present a simple fix to this problem by constructing stochastic processes that fully exploit the combinatorial structures, hence the name\n                    <jats:italic toggle=\"yes\">ComboStoc.<\/jats:italic>\n                    Using this simple strategy, we show that network training is significantly accelerated across diverse data modalities, including images and 3D structured shapes. Moreover,\n                    <jats:italic toggle=\"yes\">ComboStoc<\/jats:italic>\n                    enables a new way of test time generation which uses asynchronous time steps for different dimensions and attributes, thus allowing for varying degrees of control over them. Our code is available at: https:\/\/github.com\/Xrvitd\/ComboStoc.\n                  <\/jats:p>","DOI":"10.1145\/3811285","type":"journal-article","created":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:05:51Z","timestamp":1783062351000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["ComboStoc: Combinatorial Stochasticity for Diffusion Generative Models"],"prefix":"10.1145","volume":"45","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8273-1808","authenticated-orcid":false,"given":"Rui","family":"Xu","sequence":"first","affiliation":[{"name":"University of Hong Kong, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6049-4458","authenticated-orcid":false,"given":"Jiepeng","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Hong Kong, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3628-9777","authenticated-orcid":false,"given":"Hao","family":"Pan","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3768-6654","authenticated-orcid":false,"given":"Yang","family":"Liu","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8788-2453","authenticated-orcid":false,"given":"Xin","family":"Tong","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8452-8723","authenticated-orcid":false,"given":"Shiqing","family":"Xin","sequence":"additional","affiliation":[{"name":"Shandong University, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1231-3392","authenticated-orcid":false,"given":"Changhe","family":"Tu","sequence":"additional","affiliation":[{"name":"Shandong University, Qingdao, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2729-5860","authenticated-orcid":false,"given":"Taku","family":"Komura","sequence":"additional","affiliation":[{"name":"The University of Hong Kong, Hong Kong, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2284-3952","authenticated-orcid":false,"given":"Wenping","family":"Wang","sequence":"additional","affiliation":[{"name":"Texas A&amp;M University, Texas, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,3]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Stochastic Interpolants: A Unifying Framework for Flows and Diffusions.","author":"Albergo Michael S.","year":"2023","unstructured":"Michael S. Albergo, Nicholas M. Boffi, and Eric Vanden-Eijnden. 2023. Stochastic Interpolants: A Unifying Framework for Flows and Diffusions."},{"key":"e_1_2_1_2_1","volume-title":"Building Normalizing Flows with Stochastic Interpolants. In International Conference on Learning Representations.","author":"Albergo Michael Samuel","year":"2023","unstructured":"Michael Samuel Albergo and Eric Vanden-Eijnden. 2023. Building Normalizing Flows with Stochastic Interpolants. In International Conference on Learning Representations."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-0759"},{"key":"e_1_2_1_4_1","volume-title":"Diffdock: Diffusion steps, twists, and turns for molecular docking. arXiv preprint arXiv:2210.01776","author":"Corso Gabriele","year":"2022","unstructured":"Gabriele Corso, Hannes St\u00e4rk, Bowen Jing, Regina Barzilay, and Tommi Jaakkola. 2022. Diffdock: Diffusion steps, twists, and turns for molecular docking. arXiv preprint arXiv:2210.01776 (2022)."},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Jia Deng Wei Dong Richard Socher Li-Jia Li Kai Li and Li Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. In Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_6_1","volume-title":"International Conference on Learning Representations","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. International Conference on Learning Representations (2021)."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.02117"},{"key":"e_1_2_1_8_1","volume-title":"Addressing negative transfer in diffusion models. Advances in Neural Information Processing Systems 36","author":"Go Hyojun","year":"2024","unstructured":"Hyojun Go, Yunsung Lee, Seunghyun Lee, Shinhyeok Oh, Hyeongdon Moon, and Seungtaek Choi. 2024. Addressing negative transfer in diffusion models. Advances in Neural Information Processing Systems 36 (2024)."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00684"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00684"},{"key":"e_1_2_1_11_1","volume-title":"Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor","author":"Harris Charles R","year":"2020","unstructured":"Charles R Harris, K Jarrod Millman, St\u00e9fan J Van Der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J Smith, et al. 2020. Array programming with NumPy. nature 585, 7825 (2020), 357\u2013362."},{"key":"e_1_2_1_12_1","volume-title":"Neural Information Processing Systems","author":"Heusel Martin","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local nash equilibrium. In Neural Information Processing Systems. Curran Associates Inc."},{"key":"e_1_2_1_13_1","volume-title":"Neural Information Processing Systems","volume":"33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising Diffusion Probabilistic Models. In Neural Information Processing Systems, Vol. 33. Curran Associates, Inc."},{"key":"e_1_2_1_14_1","volume-title":"Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation. In The Fourteenth International Conference on Learning Representations.","author":"Hu Zijing","year":"2026","unstructured":"Zijing Hu, Yunze Tong, Fengda Zhang, Junkun Yuan, Jun Xiao, and Kun Kuang. 2026. Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation. In The Fourteenth International Conference on Learning Representations."},{"key":"e_1_2_1_15_1","unstructured":"Jialei Huang Guanqi Zhan Qingnan Fan Kaichun Mo Lin Shao Baoquan Chen Leonidas Guibas and Hao Dong. 2020. Generative 3D Part Assembly via Dynamic Graph Learning. In Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-2853"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.70040"},{"key":"e_1_2_1_18_1","volume-title":"Back to basics: Let denoising generative models denoise. arXiv preprint arXiv:2511.13720","author":"Li Tianhong","year":"2025","unstructured":"Tianhong Li and Kaiming He. 2025. Back to basics: Let denoising generative models denoise. arXiv preprint arXiv:2511.13720 (2025)."},{"key":"e_1_2_1_19_1","volume-title":"Flow Matching for Generative Modeling. In International Conference on Learning Representations.","author":"Lipman Yaron","year":"2023","unstructured":"Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. 2023. Flow Matching for Generative Modeling. In International Conference on Learning Representations."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3641519.3657482"},{"key":"e_1_2_1_21_1","volume-title":"International Conference on Learning Representations.","author":"Liu Xingchao","year":"2023","unstructured":"Xingchao Liu, Chengyue Gong, and Qiang Liu. 2023. Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow. In International Conference on Learning Representations."},{"key":"e_1_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Andreas Lugmayr Martin Danelljan Andres Romero Fisher Yu Radu Timofte and Luc Van Gool. 2022. RePaint: Inpainting Using Denoising Diffusion Probabilistic Models. In Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR52688.2022.01117"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-72980-5_2"},{"key":"e_1_2_1_24_1","volume-title":"Sdedit: Guided image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073","author":"Meng Chenlin","year":"2021","unstructured":"Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. 2021. Sdedit: Guided image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073 (2021)."},{"key":"e_1_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Chenlin Meng Robin Rombach Ruiqi Gao Diederik Kingma Stefano Ermon Jonathan Ho and Tim Salimans. 2023. On Distillation of Guided Diffusion Models. In Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR52729.2023.01374"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3355089.3356527"},{"key":"e_1_2_1_27_1","unstructured":"Kaichun Mo Shilin Zhu Angel X. Chang Li Yi Subarna Tripathi Leonidas J. Guibas and Hao Su. 2019b. PartNet: A Large-Scale Benchmark for Fine-Grained and Hierarchical Part-Level 3D Object Understanding. In Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_1_28_1","volume-title":"Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32 (2019)."},{"key":"e_1_2_1_29_1","volume-title":"Daniel Cohen-Or, and Dani Lischinski.","author":"Pearl Naama","year":"2023","unstructured":"Naama Pearl, Yaron Brodsky, Dana Berman, Assaf Zomet, Alex Rav Acha, Daniel Cohen-Or, and Dani Lischinski. 2023. Svnr: Spatially-variant noise removal with denoising diffusion. arXiv preprint arXiv:2306.16052 (2023)."},{"key":"e_1_2_1_30_1","volume-title":"Scalable Diffusion Models with Transformers. In International Conference on Computer Vision (ICCV).","author":"Peebles William","year":"2023","unstructured":"William Peebles and Saining Xie. 2023. Scalable Diffusion Models with Transformers. In International Conference on Computer Vision (ICCV)."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3406703"},{"key":"e_1_2_1_32_1","doi-asserted-by":"crossref","unstructured":"Robin Rombach Andreas Blattmann Dominik Lorenz Patrick Esser and Bj\u00f6rn Ommer. 2022. High-Resolution Image Synthesis With Latent Diffusion Models. In Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_2_1_33_1","unstructured":"David Ruhe Jonathan Heek Tim Salimans and Emiel Hoogeboom. 2024. Rolling Diffusion Models. arXiv:2402.09470 [cs.LG] https:\/\/arxiv.org\/abs\/2402.09470"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3354"},{"key":"e_1_2_1_35_1","volume-title":"Deeply supervised flow-based generative models. arXiv preprint arXiv:2503.14494","author":"Shin Inkyu","year":"2025","unstructured":"Inkyu Shin, Chenglin Yang, and Liang-Chieh Chen. 2025. Deeply supervised flow-based generative models. arXiv preprint arXiv:2503.14494 (2025)."},{"key":"e_1_2_1_36_1","volume-title":"History-guided video diffusion. arXiv preprint arXiv:2502.06764","author":"Song Kiwhan","year":"2025","unstructured":"Kiwhan Song, Boyuan Chen, Max Simchowitz, Yilun Du, Russ Tedrake, and Vincent Sitzmann. 2025. History-guided video diffusion. arXiv preprint arXiv:2502.06764 (2025)."},{"key":"e_1_2_1_37_1","volume-title":"Consistency Models. In International Conference on Machine Learning.","author":"Song Yang","year":"2023","unstructured":"Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. 2023. Consistency Models. In International Conference on Machine Learning."},{"key":"e_1_2_1_38_1","volume-title":"International Conference on Learning Representations.","author":"Song Yang","year":"2021","unstructured":"Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2021. Score-Based Generative Modeling through Stochastic Differential Equations. In International Conference on Learning Representations."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00690"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818094"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3732934"},{"key":"e_1_2_1_43_1","volume-title":"A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training. arXiv preprint arXiv:2405.17403","author":"Wang Kai","year":"2024","unstructured":"Kai Wang, Mingjia Shi, Yukun Zhou, Zekai Li, Zhihang Yuan, Yuzhang Shang, Xiaojiang Peng, Hanwang Zhang, and Yang You. 2024. A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training. arXiv preprint arXiv:2405.17403 (2024)."},{"key":"e_1_2_1_44_1","volume-title":"Protein structure generation via folding diffusion. Nature communications 15, 1","author":"Wu Kevin E","year":"2024","unstructured":"Kevin E Wu, Kevin K Yang, Rianne van den Berg, Sarah Alamdari, James Y Zou, Alex X Lu, and Ava P Amini. 2024. Protein structure generation via folding diffusion. Nature communications 15, 1 (2024), 1059."},{"key":"e_1_2_1_45_1","volume-title":"Structured 3d latents for scalable and versatile 3d generation. arXiv preprint arXiv:2412.01506","author":"Xiang Jianfeng","year":"2024","unstructured":"Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong Chen, Xin Tong, and Jiaolong Yang. 2024. Structured 3d latents for scalable and versatile 3d generation. arXiv preprint arXiv:2412.01506 (2024)."},{"key":"e_1_2_1_46_1","unstructured":"Xinhao Yan Jiachen Xu Yang Li Changfeng Ma Yunhan Yang Chunshi Wang Zibo Zhao Zeqiang Lai Yunfei Zhao Zhuo Chen et al. 2025. X-part: high fidelity and structure coherent shape decomposition. arXiv preprint arXiv:2509.08643 (2025)."},{"key":"e_1_2_1_47_1","volume-title":"Diffusion Models: A Comprehensive Survey of Methods and Applications. ACM Comput. Surv. 56, 4, Article 105 (Nov.","author":"Yang Ling","year":"2023","unstructured":"Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2023. Diffusion Models: A Comprehensive Survey of Methods and Applications. ACM Comput. Surv. 56, 4, Article 105 (Nov. 2023), 39 pages."},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1002\/wcms.1711"},{"key":"e_1_2_1_49_1","volume-title":"Representation alignment for generation: Training diffusion transformers is easier than you think. arXiv preprint arXiv:2410.06940","author":"Yu Sihyun","year":"2024","unstructured":"Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. 2024. Representation alignment for generation: Training diffusion transformers is easier than you think. arXiv preprint arXiv:2410.06940 (2024)."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3592442"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3658146"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3730840"},{"key":"e_1_2_1_53_1","unstructured":"Zibo Zhao Zeqiang Lai Qingxiang Lin Yunfei Zhao Haolin Liu Shuhui Yang Yifei Feng Mingxin Yang Sheng Zhang Xianghui Yang et al. 2025. Hunyuan3d 2.0: Scaling diffusion models for high resolution textured 3d assets generation. arXiv preprint arXiv:2501.12202 (2025)."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-72646-0_7"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3592103"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","deposited":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:39:45Z","timestamp":1783064385000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811285"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,3]]},"references-count":55,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,7,3]]}},"alternative-id":["10.1145\/3811285"],"URL":"https:\/\/doi.org\/10.1145\/3811285","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,3]]},"assertion":[{"value":"2026-01-07","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}