{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T04:02:06Z","timestamp":1784520126062,"version":"3.55.0"},"reference-count":29,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T00:00:00Z","timestamp":1784505600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T00:00:00Z","timestamp":1784505600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Vis. Comput. Ind. Biomed. Art"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Text-to-motion generation aims to synthesize semantically consistent and naturally coherent motion sequences from natural language descriptions. Given the continuous nature of human motion, diffusion models operating in a continuous latent space offer inherent advantages over vector quantization-based methods, particularly in avoiding quantization errors and in modeling quality. However, existing diffusion models primarily rely on mean squared error loss. This stepwise regression paradigm often leads to \u2018over-smoothed\u2019 motion sequences and struggles to capture the subtle semantic nuances embedded in textual descriptions. To realize the potential for continuous diffusion generation, an enhanced latent-space diffusion framework designed to elevate generation capabilities across two dimensions, namely, distribution approximation and semantic alignment, is proposed. Specifically, a latent-space adversarial discriminator is incorporated. By applying decoupled adversarial supervision, this component mitigates the detail loss caused by mean regression, significantly enhancing the physical realism and dynamic sharpness. Concurrently, a latent-space contrastive alignment strategy is introduced during the denoising process that reinforces the correspondence of the generated motion sequences with the given textual inputs via explicit cross-modal constraints. Extensive experiments on standard benchmarks demonstrate that the proposed method effectively addresses the limitations of conventional diffusion models, thus validating the potential of continuous diffusion frameworks within the domain of text-driven motion synthesis.<\/jats:p>","DOI":"10.1186\/s42492-026-00224-2","type":"journal-article","created":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T03:07:58Z","timestamp":1784516878000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Aligned and realistic latent diffusion for text-to-motion generation"],"prefix":"10.1186","volume":"9","author":[{"given":"Zhaowu","family":"Li","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5200-640X","authenticated-orcid":false,"given":"Rui","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Deheng","family":"Zhu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dongsheng","family":"Zhou","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaopeng","family":"Wei","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,7,20]]},"reference":[{"key":"224_CR1","doi-asserted-by":"publisher","unstructured":"Guo C, Zou SH, Zuo XX, Wang S, Ji W, Li XY et al (2022) Generating diverse and natural 3D human motions from text. In: Proceedings of the 2022 IEEE\/CVF conference on computer vision and pattern recognition, IEEE, New Orleans, 18\u201324 June 2022. https:\/\/doi.org\/10.1109\/CVPR52688.2022.00509","DOI":"10.1109\/CVPR52688.2022.00509"},{"key":"224_CR2","doi-asserted-by":"publisher","unstructured":"Guo C, Zuo XX, Wang S, Cheng L (2022) TM2T: Stochastic and tokenized modeling for the reciprocal generation of 3D human motions and texts. In: Avidan S, Brostow G, Ciss\u00e9 M, Farinella GM, Hassner T (eds) 17th European conference on computer vision. Springer, Tel Aviv, Israel, pp 580\u2013597. https:\/\/doi.org\/10.1007\/978-3-031-19833-5_34","DOI":"10.1007\/978-3-031-19833-5_34"},{"key":"224_CR3","doi-asserted-by":"publisher","unstructured":"Tevet G, Gordon B, Hertz A, Bermano AH, Cohen-Or D (2022) MotionCLIP: Exposing human motion generation to CLIP space. In: Avidan S, Brostow G, Ciss\u00e9 M, Farinella GM, Hassner T (eds) 17th European conference on computer vision. Springer, Tel Aviv, Israel, pp 358\u2013374. https:\/\/doi.org\/10.1007\/978-3-031-20047-2_21","DOI":"10.1007\/978-3-031-20047-2_21"},{"key":"224_CR4","unstructured":"Tevet G, Raab S, Gordon B, Shafir Y, Cohen-Or D, Bermano AH (2022) Human motion diffusion model. arXiv preprint arXiv:2209.14916"},{"issue":"6","key":"224_CR5","doi-asserted-by":"publisher","first-page":"4115","DOI":"10.1109\/TPAMI.2024.3355414","volume":"46","author":"MY Zhang","year":"2024","unstructured":"Zhang MY, Cai Z, Pan L, Hong FZ, Guo XY, Yang L et al (2024) MotionDiffuse: Text-driven human motion generation with diffusion model. IEEE Trans Pattern Anal Mach Intell 46(6):4115\u20134128. https:\/\/doi.org\/10.1109\/TPAMI.2024.3355414","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"224_CR6","doi-asserted-by":"publisher","unstructured":"Zhang MY, Guo XY, Pan L, Cai Z, Hong FZ, Li HR et al (2023) ReMoDiffuse: Retrieval-augmented motion diffusion model. In: Proceedings of the 2023 IEEE\/CVF international conference on computer vision, IEEE, Paris, 1\u20136 October 2023. https:\/\/doi.org\/10.1109\/ICCV51070.2023.00040","DOI":"10.1109\/ICCV51070.2023.00040"},{"key":"224_CR7","doi-asserted-by":"publisher","unstructured":"Chen X, Jiang B, Liu W, Huang ZL, Fu B, Chen T et al (2023) Executing your commands via motion diffusion in latent space. In: Proceedings of the 2023 IEEE\/CVF conference on computer vision and pattern recognition, IEEE, Vancouver, 17\u201324 June 2023. https:\/\/doi.org\/10.1109\/CVPR52729.2023.01726","DOI":"10.1109\/CVPR52729.2023.01726"},{"key":"224_CR8","doi-asserted-by":"publisher","unstructured":"Zhang ZY, Liu A, Reid I, Hartley R, Zhuang BH, Tang H (2025) Motion mamba: Efficient and long sequence motion generation. In: Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T, Varol G (eds) 18th European conference on computer vision. Springer, Milan, Italy, pp 265\u2013282. https:\/\/doi.org\/10.1007\/978-3-031-73232-4_15","DOI":"10.1007\/978-3-031-73232-4_15"},{"key":"224_CR9","doi-asserted-by":"publisher","unstructured":"Meng ZC, Xie YM, Peng XG, Han ZY, Jiang HZ (2025) Rethinking diffusion for text-driven human motion generation: Redundant representations, evaluation, and masked autoregression. In: Proceedings of the 2025 IEEE\/CVF conference on computer vision and pattern recognition, IEEE, Nashville, 10\u201317 June 2025. https:\/\/doi.org\/10.1109\/CVPR52734.2025.02594","DOI":"10.1109\/CVPR52734.2025.02594"},{"key":"224_CR10","unstructured":"Chen XY (2024) Text-driven human motion generation with motion masked diffusion model. arXiv preprint arXiv:2409.19686"},{"key":"224_CR11","unstructured":"Wang ZD, Zheng HJ, He PC, Chen WZ, Zhou MY (2023) Diffusion-GAN: Training GANs with diffusion. arXiv preprint arXiv:2206.02262"},{"key":"224_CR12","doi-asserted-by":"publisher","unstructured":"Kang M, Zhang R, Barnes C, Paris S, Kwak S, Park J et al (2025) Distilling diffusion models into conditional GANs. In: Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T, Varol G (eds) 18th European conference on computer vision. Springer, Milan, Italy, pp 428\u2013447. https:\/\/doi.org\/10.1007\/978-3-031-73390-1_25","DOI":"10.1007\/978-3-031-73390-1_25"},{"key":"224_CR13","doi-asserted-by":"publisher","unstructured":"Xu ZC, Zhang JF, Liew JH, Yan HS, Liu JW, Zhang CX et al (2024) MagicAnimate: Temporally consistent human image animation using diffusion model. In: Proceedings of the 2024 IEEE\/CVF conference on computer vision and pattern recognition, IEEE, Seattle, 16\u201322 June 2024. https:\/\/doi.org\/10.1109\/CVPR52733.2024.00147","DOI":"10.1109\/CVPR52733.2024.00147"},{"key":"224_CR14","doi-asserted-by":"publisher","unstructured":"Li JF, Cao JK, Zhang HT, Rempe D, Kautz J, Iqbal U et al (2025) GENMO: A generalist model for human motion. arXiv preprint arXiv:2505.01425. https:\/\/doi.org\/10.1109\/ICCV51701.2025.01094","DOI":"10.1109\/ICCV51701.2025.01094"},{"key":"224_CR15","doi-asserted-by":"publisher","unstructured":"Ahn H, Ha T, Choi Y, Yoo H, Oh S (2018) Text2Action: Generative adversarial synthesis from language to action. In: Proceedings of the 2018 IEEE international conference on robotics and automation (ICRA), IEEE, Brisbane, 21\u201325 May 2018. https:\/\/doi.org\/10.1109\/ICRA.2018.8460608","DOI":"10.1109\/ICRA.2018.8460608"},{"key":"224_CR16","doi-asserted-by":"publisher","unstructured":"Raab S, Leibovitch I, Li PZ, Aberman K, Sorkine-Hornung O, Cohen-Or D (2023) MoDi: Unconditional motion synthesis from diverse data. In: Proceedings of the 2023 IEEE\/CVF conference on computer vision and pattern recognition, IEEE, Vancouver, 17\u201324 June 2023. https:\/\/doi.org\/10.1109\/CVPR52729.2023.01333","DOI":"10.1109\/CVPR52729.2023.01333"},{"key":"224_CR17","doi-asserted-by":"publisher","unstructured":"Amballa A, Akkinapalli G, Muralikrishnan V (2025) LS-GAN: Human motion synthesis with latent-space GANs. In: Proceedings of the 2025 IEEE\/CVF winter conference on applications of computer vision workshops, IEEE, Tucson, 28 February\u20134 March 2025. https:\/\/doi.org\/10.1109\/WACVW65960.2025.00039","DOI":"10.1109\/WACVW65960.2025.00039"},{"key":"224_CR18","doi-asserted-by":"publisher","unstructured":"Petrovich M, Black MJ, Varol G (2023) TMR: Text-to-motion retrieval using contrastive 3D human motion synthesis. In: Proceedings of the 2023 IEEE\/CVF international conference on computer vision, IEEE, Paris, 1\u20136 October 2023. https:\/\/doi.org\/10.1109\/ICCV51070.2023.00870","DOI":"10.1109\/ICCV51070.2023.00870"},{"key":"224_CR19","doi-asserted-by":"publisher","unstructured":"Petrovich M, Black MJ, Varol G (2022) TEMOS: generating diverse human motions from textual descriptions. In: Avidan S, Brostow G, Ciss\u00e9 M, Farinella GM, Hassner T (eds) 17th European conference on computer vision. Springer, Tel Aviv, Israel, pp 480\u2013497. https:\/\/doi.org\/10.1007\/978-3-031-20047-2_28","DOI":"10.1007\/978-3-031-20047-2_28"},{"key":"224_CR20","unstructured":"Kinfu KA, Vidal R (2025) MotionBind: Multi-modal human motion alignment for retrieval, recognition, and generation. In: Proceedings of the 39th annual conference on neural information processing systems, NeurIPS, San Diego, 2\u20137 December 2025"},{"key":"224_CR21","doi-asserted-by":"publisher","unstructured":"Zhang PF, Liu PX, Garrido P, Kim H, Chaudhuri B (2025) Kinmo: Kinematic-aware human motion understanding and generation. In: Proceedings of the 2025 IEEE\/CVF international conference on computer vision, IEEE, Honolulu, 19\u201325 October 2025. https:\/\/doi.org\/10.1109\/ICCV51701.2025.01041","DOI":"10.1109\/ICCV51701.2025.01041"},{"key":"224_CR22","unstructured":"Lu SL, Chen LH, Zeng AL, Lin J, Zhang RM, Zhang L et al (2023) HumanTOMATO: Text-aligned whole-body motion generation. arXiv preprint arXiv:2310.12978"},{"key":"224_CR23","doi-asserted-by":"publisher","unstructured":"Uchida K, Shibuya T, Takida Y, Murata N, Tanke J, Takahashi S et al (2025) MoLA: Motion generation and editing with latent diffusion enhanced by adversarial training. In: Proceedings of the 2025 IEEE\/CVF conference on computer vision and pattern recognition workshops, IEEE, Nashville, 11\u201312 June 2025. https:\/\/doi.org\/10.1109\/CVPRW67362.2025.00274","DOI":"10.1109\/CVPRW67362.2025.00274"},{"key":"224_CR24","unstructured":"Takida Y, Imaizumi M, Shibuya T, Lai CH, Uesaka T, Murata N et al (2024) SAN: Inducing metrizability of GAN with discriminative normalized linear layer. arXiv preprint arXiv:2301.12811"},{"issue":"4","key":"224_CR25","doi-asserted-by":"publisher","first-page":"236","DOI":"10.1089\/big.2016.0028","volume":"4","author":"M Plappert","year":"2016","unstructured":"Plappert M, Mandery C, Asfour T (2016) The KIT motion-language dataset. Big Data 4(4):236\u2013252. https:\/\/doi.org\/10.1089\/big.2016.0028","journal-title":"Big Data"},{"key":"224_CR26","doi-asserted-by":"publisher","unstructured":"Pinyoanuntapong E, Wang P, Lee M, Chen C (2024) MMM: Generative masked motion model. In: Proceedings of the 2024 IEEE\/CVF conference on computer vision and pattern recognition, IEEE, Seattle, 16\u201322 June 2024. https:\/\/doi.org\/10.1109\/CVPR52733.2024.00153","DOI":"10.1109\/CVPR52733.2024.00153"},{"key":"224_CR27","doi-asserted-by":"publisher","unstructured":"Guo C, Mu YX, Javed MG, Wang S, Cheng L (2024) MoMask: Generative masked modeling of 3D human motions. In: Proceedings of the 2024 IEEE\/CVF conference on computer vision and pattern recognition, IEEE, Seattle, 16\u201322 June 2024. https:\/\/doi.org\/10.1109\/CVPR52733.2024.00186","DOI":"10.1109\/CVPR52733.2024.00186"},{"issue":"7","key":"224_CR28","doi-asserted-by":"publisher","first-page":"4277","DOI":"10.1007\/s11263-025-02392-9","volume":"133","author":"Y Wang","year":"2025","unstructured":"Wang Y, Li M, Liu JP, Leng ZY, Li FWB, Zhang ZY et al (2025) Fg-T2M++: LLMs-augmented fine-grained text driven human motion generation. Int J Comput Vis 133(7):4277\u20134293. https:\/\/doi.org\/10.1007\/s11263-025-02392-9","journal-title":"Int J Comput Vis"},{"key":"224_CR29","doi-asserted-by":"publisher","unstructured":"Pinyoanuntapong E, Saleem MU, Wang P, Lee M, Das S, Chen C (2025) BAMM: Bidirectional autoregressive motion model. In: Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T, Varol G (eds) 18th European conference on computer vision. Springer, Milan, Italy, pp 172\u2013190. https:\/\/doi.org\/10.1007\/978-3-031-72633-0_10","DOI":"10.1007\/978-3-031-72633-0_10"}],"container-title":["Visual Computing for Industry, Biomedicine, and Art"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s42492-026-00224-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s42492-026-00224-2","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s42492-026-00224-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T03:08:05Z","timestamp":1784516885000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1186\/s42492-026-00224-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,20]]},"references-count":29,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,12]]}},"alternative-id":["224"],"URL":"https:\/\/doi.org\/10.1186\/s42492-026-00224-2","relation":{},"ISSN":["2524-4442"],"issn-type":[{"value":"2524-4442","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,20]]},"assertion":[{"value":"4 February 2026","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 June 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 July 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare that they have no competing interests.","order":1,"name":"Ethics","label":"Competing interests","group":{"name":"EthicsHeading","label":"Declarations"}}],"article-number":"14"}}