{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:22:19Z","timestamp":1783063339969,"version":"3.54.6"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T00:00:00Z","timestamp":1783036800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"The National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62372336"],"award-info":[{"award-number":["62372336"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"The Key Research and Development Program for Technological Innovation Project of Hubei Province","award":["2025BAB020"],"award-info":[{"award-number":["2025BAB020"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2026,7,3]]},"abstract":"<jats:p>We present HumanFlow, a unified flow-matching-based framework that enables high-fidelity and controllable full-body human image generation under diverse human-centric control conditions. Despite recent progress, controllable human image generation poses a fundamental challenge in balancing high visual fidelity with strict adherence to human-centric control conditions. HumanFlow formulates human image generation as a conditional flow-matching process with deterministic generation dynamics. To incorporate such human-centric control conditions into the pretrained model, we introduce a unified control framework with Control Encoder and Token-ControlNet. A Control Encoder maps diverse conditions into a unified latent representation that is spatially aligned with the image latent space. Token-ControlNet is a lightweight control network architecturally aligned with the FLUX double-stream design. To address accurate structural control over human bodies, we further propose the Human Topology Consistency Loss (HTCL). HTCL regularizes conditional flow matching by constraining generated human configurations to a union of statistically grounded topology manifolds defined by normalized bone ratios and joint angles. To support large-scale training and systematic evaluation, we construct MiCoGen, a multi-condition human image dataset comprising over one million full-body human images with aligned text descriptions and rich human-centric control conditions. Extensive quantitative and qualitative evaluations on the MiCoGen dataset show that HumanFlow consistently achieves improved structural consistency than the existing diffusion-based and flow-matching-based methods, while maintaining high visual fidelity.<\/jats:p>","DOI":"10.1145\/3811361","type":"journal-article","created":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:05:51Z","timestamp":1783062351000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["HumanFlow: Controllable Human Image Generation via Flow Matching"],"prefix":"10.1145","volume":"45","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-1258-5442","authenticated-orcid":false,"given":"Wenzhuo","family":"Fan","sequence":"first","affiliation":[{"name":"School of Computer Science, Wuhan University, Wuhan, HuBei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1274-3373","authenticated-orcid":false,"given":"Hongsheng","family":"Zheng","sequence":"additional","affiliation":[{"name":"School of Computer Science, Wuhan University, Wuhan, HuBei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2256-6294","authenticated-orcid":false,"given":"Jianchi","family":"Sun","sequence":"additional","affiliation":[{"name":"School of Computer Science, Wuhan University, Wuhan, HuBei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-7920-4843","authenticated-orcid":false,"given":"Fei","family":"Fang","sequence":"additional","affiliation":[{"name":"Wuhan Textile University, Wuhan, Hubei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9946-2448","authenticated-orcid":false,"given":"Hong","family":"Ding","sequence":"additional","affiliation":[{"name":"Guangxi University of Finance and Economic, Nanning, GuangXi, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4526-6297","authenticated-orcid":false,"given":"Chunxia","family":"Xiao","sequence":"additional","affiliation":[{"name":"School of Computer Science, Wuhan University, Wuhan, HuBei, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,3]]},"reference":[{"key":"e_1_2_2_1_1","volume-title":"Building normalizing flows with stochastic interpolants. arXivpreprint arXiv:2209.15571","author":"Albergo Michael S","year":"2022","unstructured":"Michael S Albergo and Eric Vanden-Eijnden. 2022. Building normalizing flows with stochastic interpolants. arXivpreprint arXiv:2209.15571 (2022)."},{"key":"e_1_2_2_2_1","volume-title":"International Conference on Machine Learning. PMLR","author":"Bao Fan","year":"2023","unstructured":"Fan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li, Shi Pu, Yaole Wang, Gang Yue, Yue Cao, Hang Su, and Jun Zhu. 2023. One transformer fits all distributions in multi-modal diffusion at scale. In International Conference on Machine Learning. PMLR, London, UK, 1692\u20131717."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3641519.3657525"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3746252.3761227"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02676"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02676"},{"key":"e_1_2_2_7_1","volume-title":"Diffusion models beat gans on image synthesis. Advances in neural information processing systems 34","author":"Dhariwal Prafulla","year":"2021","unstructured":"Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion models beat gans on image synthesis. Advances in neural information processing systems 34 (2021), 8780\u20138794."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19787-1_1"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i3.32330"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01774"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i22.34571"},{"key":"e_1_2_2_12_1","volume-title":"VITON: An Image-based Virtual Try-on Network. In CVPR.","author":"Han Xintong","year":"2018","unstructured":"Xintong Han, Zuxuan Wu, Zhe Wu, Ruichi Yu, and Larry S Davis. 2018. VITON: An Image-based Virtual Try-on Network. In CVPR. Los Alamitos, CA, USA."},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"e_1_2_2_14_1","volume-title":"Ronan Le Bras, and Yejin Choi","author":"Hessel Jack","year":"2021","unstructured":"Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021. Clipscore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718 (2021)."},{"key":"e_1_2_2_15_1","volume-title":"Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_16_1","volume-title":"Denoising diffusion probabilistic models. Advances in neural information processing systems 33","author":"Ho Jonathan","year":"2020","unstructured":"Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840\u20136851."},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.02675"},{"key":"e_1_2_2_18_1","volume-title":"From parts to whole: A unified reference framework for controllable human image generation. arXiv preprint arXiv:2404.15267","author":"Huang Zehuan","year":"2024","unstructured":"Zehuan Huang, Hongxing Fan, Lipeng Wang, and Lu Sheng. 2024. From parts to whole: A unified reference framework for controllable human image generation. arXiv preprint arXiv:2404.15267 (2024)."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01185"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3721238.3730663"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3550454.3555513"},{"key":"e_1_2_2_22_1","volume-title":"Stage-wise dynamics of classifier-free guidance in diffusion models. arXiv preprint arXiv:2509.22007","author":"Jin Cheng","year":"2025","unstructured":"Cheng Jin, Qitan Shi, and Yuantao Gu. 2025. Stage-wise dynamics of classifier-free guidance in diffusion models. arXiv preprint arXiv:2509.22007 (2025)."},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01465"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00582"},{"key":"e_1_2_2_25_1","volume-title":"Peter Selednik, Stuart Anderson, and Shunsuke Saito.","author":"Khirodkar Rawal","year":"2024","unstructured":"Rawal Khirodkar, Timur Bagautdinov, Julieta Martinez, Su Zhaoen, Austin James, Peter Selednik, Stuart Anderson, and Shunsuke Saito. 2024. Sapiens: Foundation for Human Vision Models. arXiv preprint arXiv:2408.12569 (2024)."},{"key":"e_1_2_2_26_1","doi-asserted-by":"crossref","unstructured":"Taekyung Ki Dongchan Min and Gyeongsu Chae. 2025. Float: Generative motion latent flow matching for audio-driven talking portrait. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 14699\u201314710. Jeongho Kim Guojung Gu Minho Park Sunghyun Park and Jaegul Choo. 2024. Stableviton: Learning semantic correspondence with latent diffusion model for virtual try-on. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. Los Alamitos CA USA 8176\u20138185.","DOI":"10.1109\/ICCV51701.2025.01364"},{"key":"e_1_2_2_27_1","unstructured":"Black Forest Labs. 2024. FLUX. https:\/\/github.com\/black-forest-labs\/flux."},{"key":"e_1_2_2_28_1","volume-title":"Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space. arXiv preprint arXiv:2506.15742","author":"Labs Black Forest","year":"2025","unstructured":"Black Forest Labs, Stephen Batifol, Andreas Blattmann, Frederic Boesel, Saksham Consul, Cyril Diagne, Tim Dockhorn, Jack English, Zion English, Patrick Esser, et al. 2025. FLUX. 1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space. arXiv preprint arXiv:2506.15742 (2025)."},{"key":"e_1_2_2_29_1","unstructured":"Ming Li Taojiannan Yang Huafeng Kuang Jie Wu Zhaoning Wang Xuefeng Xiao and Chen Chen. 2024b. ControlNet ++: Improving Conditional Controls with"},{"key":"e_1_2_2_30_1","volume-title":"Consistency Feedback. In European Conference on Computer Vision","author":"Efficient","unstructured":"Efficient Consistency Feedback. In European Conference on Computer Vision. Cham, Switzerland."},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00664"},{"key":"e_1_2_2_32_1","volume-title":"Flow Matching for Generative Modeling. In The Eleventh International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=PqvMRDCJT9t","author":"Lipman Yaron","year":"2023","unstructured":"Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. 2023. Flow Matching for Generative Modeling. In The Eleventh International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=PqvMRDCJT9t"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00263"},{"key":"e_1_2_2_34_1","volume-title":"The Eleventh International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=XVjTT1nw5z","author":"Liu Xingchao","year":"2023","unstructured":"Xingchao Liu, Chengyue Gong, and Qiang Liu. 2023. Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow. In The Eleventh International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=XVjTT1nw5z"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i4.28189"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i5.28226"},{"key":"e_1_2_2_37_1","unstructured":"OpenAI. 2024. DALL-E 3: A Powerful AI for Creating Images from Text. https:\/\/openai.com\/dall-e-3"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3757377.3763984"},{"key":"e_1_2_2_39_1","volume-title":"Unicontrol: A unified diffusion model for controllable visual generation in the wild","author":"Qin C.","year":"2023","unstructured":"C. Qin, S. Zhang, N. Yu, Y. Feng, X. Yang, Y. Zhou, H. Wang, J.C. Niebles, C. Xiong, S. Savarese, et al. 2023. Unicontrol: A unified diffusion model for controllable visual generation in the wild. In NeurIPS. Cambridge, MA, USA."},{"key":"e_1_2_2_40_1","volume-title":"International conference on machine learning. PMLR","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PMLR, London, UK, 8748\u20138763."},{"key":"e_1_2_2_41_1","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21, 140 (2020), 1\u201367.","journal-title":"Journal of machine learning research"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"e_1_2_2_44_1","unstructured":"J. Song C. Meng and S. Ermon. 2021. Denoising diffusion implicit models. In ICLR. Redondo Beach CA USA."},{"key":"e_1_2_2_45_1","volume-title":"Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456","author":"Song Yang","year":"2020","unstructured":"Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2020. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456 (2020)."},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v40i11.37875"},{"key":"e_1_2_2_47_1","volume-title":"An improved fire detection approach based on YOLO-v8 for smart cities. Neural computing and applications 35, 28","author":"Talaat Fatma M","year":"2023","unstructured":"Fatma M Talaat and Hanaa ZainEldin. 2023. An improved fire detection approach based on YOLO-v8 for smart cities. Neural computing and applications 35, 28 (2023), 20939\u201320954."},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01386"},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3592451"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i7.32827"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00807"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.52202\/079017-4068"},{"key":"e_1_2_2_53_1","volume-title":"Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis. arXiv preprint arXiv:2306.09341","author":"Wu Xiaoshi","year":"2023","unstructured":"Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li. 2023. Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis. arXiv preprint arXiv:2306.09341 (2023)."},{"key":"e_1_2_2_54_1","unstructured":"Yuxin Wu Alexander Kirillov Francisco Massa Wan-Yen Lo and Ross Girshick. 2019. Detectron2. https:\/\/github.com\/facebookresearch\/detectron2."},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00157"},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i9.32973"},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3757377.3763815"},{"key":"e_1_2_2_58_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Yuanzhen Li Chunxia Xiao","year":"2024","unstructured":"Chunxia Xiao Yuanzhen Li, Fei Luo. 2024. Diffusion-FOF: Single-view Clothed Human Reconstruction via Diffusion-based Fourier Occupancy Field. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"e_1_2_2_60_1","volume-title":"Flashvideo: Flowing fidelity to detail for efficient high-resolution video generation. arXivpreprint arXiv:2502.05179","author":"Zhang Shilong","year":"2025","unstructured":"Shilong Zhang, Wenbo Li, Shoufa Chen, Chongjian Ge, Peize Sun, Yida Zhang, Yi Jiang, Zehuan Yuan, Binyue Peng, and Ping Luo. 2025a. Flashvideo: Flowing fidelity to detail for efficient high-resolution video generation. arXivpreprint arXiv:2502.05179 (2025)."},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51701.2025.01814"},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i10.33146"},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3471875"},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00238"},{"key":"e_1_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00010"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","deposited":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T07:13:12Z","timestamp":1783062792000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3811361"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,3]]},"references-count":65,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,7,3]]}},"alternative-id":["10.1145\/3811361"],"URL":"https:\/\/doi.org\/10.1145\/3811361","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,3]]},"assertion":[{"value":"2026-01-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}