{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T11:40:59Z","timestamp":1786534859917,"version":"3.56.0"},"reference-count":44,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T00:00:00Z","timestamp":1786492800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T00:00:00Z","timestamp":1786492800000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>We introduce stylized phase manifolds\u2014a compact, interpretable latent representation that disentangles motion content (e.g. \u201cjumping\u201d, \u201cwalking\u201d), the temporal structure (e.g. motion cycle frequency, gait timing), and style (i.e. how the motion is performed). Learned in an unsupervised manner and inherently low\u2010dimensional, the manifold offers intuitive and flexible editing. Building on this representation, we develop a diffusion\u2010based motion generator that enables fine\u2010grained control over semantic, temporal, and stylistic aspects of motion. To connect high\u2010level intent with low\u2010level motion, we treat the stylized manifold as an intermediate representation\u2014a structured bridge between natural language and motion. By first mapping text into this manifold, our two\u2010stage pipeline improves the control over for text\u2010based motion generation, while producing high\u2010quality, diverse motion outputs.<\/jats:p>","DOI":"10.1111\/cgf.70563","type":"journal-article","created":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T11:18:02Z","timestamp":1786533482000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["MotionPyramid: Controllable Motion Synthesis via Stylized Phase Manifolds"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7963-5014","authenticated-orcid":false,"given":"Jingyuan","family":"Li","sequence":"first","affiliation":[{"name":"ETH Zurich  Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9309-9967","authenticated-orcid":false,"given":"Peizhuo","family":"Li","sequence":"additional","affiliation":[{"name":"ETH Zurich  Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7754-0791","authenticated-orcid":false,"given":"Andreas","family":"Aristidou","sequence":"additional","affiliation":[{"name":"University of Cyprus  Cyprus"},{"name":"CYENS Centre of Excellence  Cyprus"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8089-3974","authenticated-orcid":false,"given":"Olga","family":"Sorkine\u2010Hornung","sequence":"additional","affiliation":[{"name":"ETH Zurich  Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2026,8,12]]},"reference":[{"key":"e_1_2_7_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275038"},{"key":"e_1_2_7_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/566654.566606"},{"key":"e_1_2_7_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392469"},{"key":"e_1_2_7_5_2","doi-asserted-by":"crossref","unstructured":"Barquero German Escalera Sergio andPalmero Cristina. \u201cBelfusion: Latent diffusion for behavior\u2010driven human motion prediction\u201d.Proceedings of the IEEE\/CVF International Conference on Computer Vision.2023 2317\u201323273.","DOI":"10.1109\/ICCV51070.2023.00220"},{"key":"e_1_2_7_6_2","doi-asserted-by":"crossref","unstructured":"Barquero German Escalera Sergio andPalmero Cristina. \u201cSeamless human motion composition with blended positional encodings\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 457\u20134693 9 10.","DOI":"10.1109\/CVPR52733.2024.00051"},{"key":"e_1_2_7_7_2","doi-asserted-by":"crossref","unstructured":"Chen Xin Jiang Biao Liu Wen et al. \u201cExecuting your commands via motion diffusion in latent space\u201d.Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition.2023 18000\u2013180103.","DOI":"10.1109\/CVPR52729.2023.01726"},{"key":"e_1_2_7_8_2","first-page":"8780","article-title":"Diffusion models beat gans on image synthesis","volume":"34","author":"Dhariwal Prafulla","year":"2021","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_7_9_2","doi-asserted-by":"crossref","unstructured":"Feng Yao Lin Jing Dwivedi Sai Kumar et al. \u201cChatpose: Chatting about 3d human pose\u201d.Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition.2024 2093\u201321032.","DOI":"10.1109\/CVPR52733.2024.00204"},{"key":"e_1_2_7_10_2","unstructured":"Gatys Leon A Ecker Alexander S andBethge Matthias. \u201cA neural algorithm of artistic style\u201d.arXiv preprint arXiv:1508.06576(2015) 4."},{"key":"e_1_2_7_11_2","doi-asserted-by":"crossref","unstructured":"Guo Chuan Mu Yuxuan Javed Muhammad Gohar et al. \u201cMomask: Generative masked modeling of 3d human motions\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 1900\u201319109.","DOI":"10.1109\/CVPR52733.2024.00186"},{"key":"e_1_2_7_12_2","doi-asserted-by":"crossref","unstructured":"Geng Zigang Wang Chunyu Wei Yixuan et al. \u201cHuman Pose as Compositional Tokens\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2023 660\u20136712.","DOI":"10.1109\/CVPR52729.2023.00071"},{"key":"e_1_2_7_13_2","doi-asserted-by":"crossref","unstructured":"Guo Chuan Zou Shihao Zuo Xinxin et al. \u201cGenerating Diverse and Natural 3D Human Motions From Text\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June2022 5152\u201351616.","DOI":"10.1109\/CVPR52688.2022.00509"},{"key":"e_1_2_7_14_2","doi-asserted-by":"crossref","unstructured":"Guo Chuan Zou Shihao Zuo Xinxin et al. \u201cGenerating diverse and natural 3d human motions from text\u201d.Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition.2022 5152\u201351612 8.","DOI":"10.1109\/CVPR52688.2022.00509"},{"key":"e_1_2_7_15_2","first-page":"6840","article-title":"Denoising diffusion probabilistic models","volume":"33","author":"Ho Jonathan","year":"2020","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_7_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073663"},{"key":"e_1_2_7_17_2","doi-asserted-by":"crossref","unstructured":"Huang Yiheng Yang Hui Luo Chuanchen et al. \u201cStablemofusion: Towards robust and efficient diffusion\u2010based motion generation framework\u201d.Proceedings of the 32nd ACM International Conference on Multimedia.2024 224\u20132323 12.","DOI":"10.1145\/3664647.3681657"},{"key":"e_1_2_7_18_2","first-page":"8","volume-title":"Computer Graphics Forum","author":"Hu Lei","year":"2024"},{"key":"e_1_2_7_19_2","unstructured":"Jiang Yu Chen Yixing andLi Xingyang. \u201cKETA: Kinematic\u2010Phrases\u2010Enhanced Text\u2010to\u2010Motion Generation via Finegrained Alignment\u201d.arXiv preprint arXiv:2501.15058(2025) 3."},{"key":"e_1_2_7_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3516429"},{"key":"e_1_2_7_21_2","doi-asserted-by":"crossref","unstructured":"Jang Deok\u2010Kyeong Ye Yuting Won Jungdam andLee Sung\u2010Hee. \u201cMOCHA: Real\u2010Time Motion Characterization via Context Matching\u201d.SIGGRAPH Asia 2023 Conference Papers.2023 1\u2013113.","DOI":"10.1145\/3610548.3618252"},{"key":"e_1_2_7_22_2","unstructured":"Kingma Diederik PandBa Jimmy. \u201cAdam: A method for stochastic optimization\u201d.arXiv preprint arXiv:1412.6980(2014) 6."},{"key":"e_1_2_7_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/566570.566605"},{"key":"e_1_2_7_24_2","doi-asserted-by":"crossref","unstructured":"Karunratanakul Korrawe Preechakul Konpat Aksan Emre et al. \u201cOptimizing diffusion noise can serve as universal motion priors\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2024 1334\u201313453.","DOI":"10.1109\/CVPR52733.2024.00133"},{"key":"e_1_2_7_25_2","doi-asserted-by":"crossref","unstructured":"Karunratanakul Korrawe Preechakul Konpat Suwajanakorn Supasorn andTang Siyu. \u201cGuided motion diffusion for controllable human motion synthesis\u201d.Proceedings of the IEEE\/CVF International Conference on Computer Vision.2023 2151\u201321623 12.","DOI":"10.1109\/ICCV51070.2023.00205"},{"key":"e_1_2_7_26_2","unstructured":"Li Chuqiao Chibane Julian He Yannan et al. \u201cUnimotion: Unifying 3D Human Motion Synthesis and Understanding\u201d.International Conference on 3D Vision (3DV). Mar.20253 9 10."},{"key":"e_1_2_7_27_2","first-page":"223","volume-title":"European Conference on Computer Vision","author":"Liu Xinpeng","year":"2024"},{"key":"e_1_2_7_28_2","doi-asserted-by":"crossref","unstructured":"Li Peizhuo Starke Sebastian Ye Yuting andSorkine\u2010Hornung Olga. \u201cWalkTheDog: Cross\u2010Morphology Motion Alignment via Phase Manifolds\u201d.ACM SIGGRAPH 2024 Conference Papers.2024 1\u2013101\u20134 7 12.","DOI":"10.1145\/3641519.3657508"},{"key":"e_1_2_7_29_2","doi-asserted-by":"crossref","unstructured":"Lee Yongjoon Wampler Kevin Bernstein Gilbert et al. \u201cMotion fields for interactive character locomotion\u201d.2010 1\u201382.","DOI":"10.1145\/1882262.1866160"},{"key":"e_1_2_7_30_2","first-page":"8162","volume-title":"International conference on machine learning","author":"Nichol Alexander Quinn","year":"2021"},{"key":"e_1_2_7_31_2","article-title":"PyTorch: an imperative style","volume":"12","author":"Paszke Adam","year":"2019","journal-title":"High\u2010Performance Deep Learning Library"},{"issue":"2","key":"e_1_2_7_32_2","article-title":"Hierarchical text\u2010conditional image generation with clip latents","volume":"1","author":"Ramesh Aditya","year":"2022","journal-title":"arXiv preprint arXiv:2204.06125"},{"key":"e_1_2_7_33_2","first-page":"8748","volume-title":"International conference on machine learning","author":"Radford Alec","year":"2021"},{"key":"e_1_2_7_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3407659"},{"key":"e_1_2_7_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3528223.3530178"},{"key":"e_1_2_7_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3355089.3356505"},{"key":"e_1_2_7_37_2","unstructured":"Tevet Guy Raab Sigal Gordon Brian et al. \u201cHuman Motion Diffusion Model\u201d.The Eleventh International Conference on Learning Representations.2023. url:https:\/\/openreview.net\/forum?id=SJ1kSyO2jwu2 3 6."},{"key":"e_1_2_7_38_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_7_39_2","article-title":"Neural discrete representation learning","volume":"30","author":"Van Den Oord Aaron","year":"2017","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_7_40_2","unstructured":"Xie Yiming Jampani Varun Zhong Lei et al. \u201cOmni\u2010control: Control any joint at any time for human motion generation\u201d.arXiv preprint arXiv:2310.08580(2023) 3 8 9."},{"key":"e_1_2_7_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/2766999"},{"key":"e_1_2_7_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/3658137"},{"key":"e_1_2_7_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3355414"},{"key":"e_1_2_7_44_2","first-page":"405","volume-title":"European Conference on Computer Vision","author":"Zhong Lei","year":"2024"},{"key":"e_1_2_7_45_2","doi-asserted-by":"crossref","unstructured":"Zhang Jianrong Zhang Yangsong Cun Xiaodong et al. \u201cGenerating human motion from textual descriptions with discrete representations\u201d.Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition.2023 14730\u2013147402.","DOI":"10.1109\/CVPR52729.2023.01415"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70563","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70563","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70563","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T11:18:15Z","timestamp":1786533495000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70563"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,8,12]]},"references-count":44,"alternative-id":["10.1111\/cgf.70563"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70563","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,8,12]]},"assertion":[{"value":"2026-08-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70563"}}