{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T04:47:09Z","timestamp":1777870029681,"version":"3.51.4"},"reference-count":36,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T00:00:00Z","timestamp":1777507200000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T00:00:00Z","timestamp":1777507200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Human motion data is inherently rich and complex, containing both semantic content and subtle stylistic features that are challenging to model. We propose a novel method for effective disentanglement of the style and content in human motion data to facilitate style transfer. Our approach is guided by the insight that content corresponds to coarse motion attributes while style captures the finer, expressive details. To model this hierarchy, we employ Residual Vector Quantized Variational Autoencoders (RVQ\u2010VAEs) to learn a coarse\u2010to\u2010fine representation of motion. We further enhance the disentanglement by integrating codebook learning with contrastive learning and a novel information leakage loss to organize the content and the style across different codebooks. We harness this disentangled representation using our simple and effective inference\u2010time technique\n                    <jats:italic>Quantized Code Swapping<\/jats:italic>\n                    , which enables motion style transfer without requiring any fine\u2010tuning for unseen styles. Our framework demonstrates strong versatility across multiple inference applications, including style transfer, style removal, and motion blending.\n                  <\/jats:p>","DOI":"10.1111\/cgf.70377","type":"journal-article","created":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T15:14:04Z","timestamp":1777562044000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["VQ\u2010Style: Disentangling Style and Content in Motion with Residual Quantized Representations"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-9734-2693","authenticated-orcid":false,"given":"Fatemeh","family":"Zargarbashi","sequence":"first","affiliation":[{"name":"ETH Z\u00fcrich  Switzerland"},{"name":"DisneyResearch|Studios  Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0442-9781","authenticated-orcid":false,"given":"Dhruv","family":"Agrawal","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich  Switzerland"},{"name":"DisneyResearch|Studios  Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3038-4881","authenticated-orcid":false,"given":"Jakob","family":"Buhmann","sequence":"additional","affiliation":[{"name":"DisneyResearch|Studios  Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-7496-6185","authenticated-orcid":false,"given":"Martin","family":"Guay","sequence":"additional","affiliation":[{"name":"DisneyResearch|Studios  Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6604-4784","authenticated-orcid":false,"given":"Stelian","family":"Coros","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich  Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1909-8082","authenticated-orcid":false,"given":"Robert W.","family":"Sumner","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich  Switzerland"},{"name":"DisneyResearch|Studios  Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,4,30]]},"reference":[{"key":"e_1_2_7_2_2","doi-asserted-by":"crossref","unstructured":"AgrawalD. K\u00f6nigM. BuhmannJ. SumnerR. GuayM.: Trajectory augmentation for robust neural locomotion controllers. InProceedings of the 19th International Joint Conference on Computer Vision Imaging and Computer Graphics Theory and Applications-Volume 1: GRAPP HUCAPP and IVAPP(2024) SciTePress pp.23\u201333. 6","DOI":"10.5220\/0012273300003660"},{"issue":"4","key":"e_1_2_7_3_2","doi-asserted-by":"crossref","first-page":"64","DOI":"10.1145\/3386569.3392469","article-title":"Unpaired motion style transfer from video to animation","volume":"39","author":"Aberman K.","year":"2020","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_7_4_2","unstructured":"CMU Mocap Dataset.Carnegie Mellon Universityhttps:\/\/mocap.cs.cmu.edu\/. 9"},{"issue":"7","key":"e_1_2_7_5_2","doi-asserted-by":"crossref","DOI":"10.23915\/distill.00011","article-title":"Feature-wise transformations","volume":"3","author":"Dumoulin V.","year":"2018","journal-title":"Distill"},{"key":"e_1_2_7_6_2","unstructured":"DaiM. WangJ. FanK. JiB. ZhaoH. DongJ. DaiB.: Towards synthesized and editable motion in-betweening through part-wise phase representation.arXiv preprint arXiv:2503.08180(2025). 2"},{"issue":"11","key":"e_1_2_7_7_2","doi-asserted-by":"crossref","first-page":"793","DOI":"10.1119\/1.1937609","article-title":"Transmission of information: A statistical theory of communications","volume":"29","author":"Fano R. M.","year":"1961","journal-title":"American Journal of Physics"},{"key":"e_1_2_7_8_2","unstructured":"GuoC. MuY. ZuoX. DaiP. YanY. LuJ. ChengL.: Generative human motion stylization in latent space.CoRR(2024). 2 8 9 11 13 14"},{"key":"e_1_2_7_9_2","first-page":"580","volume-title":"European Conference on Computer Vision","author":"Guo C.","year":"2022"},{"key":"e_1_2_7_10_2","unstructured":"HadjeresG. CrestelL.: Vector quantized contrastive predictive coding for template-based music generation.arXiv preprint arXiv:2004.10120(2020). 4"},{"key":"e_1_2_7_11_2","first-page":"14096","volume-title":"International Conference on Machine Learning","author":"Huh M.","year":"2023"},{"key":"e_1_2_7_12_2","unstructured":"HoJ. SalimansT.: Classifier-free diffusion guidance.arXiv preprint arXiv:2207.12598(2022). 2"},{"issue":"4","key":"e_1_2_7_13_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2897824.2925975","article-title":"A deep learning framework for character motion synthesis and editing","volume":"35","author":"Holden D.","year":"2016","journal-title":"ACM Transactions on Graphics (ToG)"},{"issue":"3","key":"e_1_2_7_14_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3516429","article-title":"Motion puzzle: Arbitrary motion style transfer by body part","volume":"41","author":"Jang D.-K.","year":"2022","journal-title":"ACM Transactions on Graphics (TOG)"},{"issue":"4","key":"e_1_2_7_15_2","doi-asserted-by":"crossref","first-page":"307","DOI":"10.1561\/2200000056","article-title":"An introduction to variational autoencoders","volume":"12","author":"Kingma D. P.","year":"2019","journal-title":"Foundations and Trends\u00ae in Machine Learning"},{"key":"e_1_2_7_16_2","first-page":"1","volume-title":"2020 International Joint Conference on Neural Networks (IJCNN)","author":"\u0141a\u0144cucki A.","year":"2020"},{"key":"e_1_2_7_17_2","unstructured":"LeeD. KimC. KimS. ChoM. HanW.-S.: Autore-gressive image generation using residual quantization. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2022) pp.11523\u201311532. 2 3 13"},{"issue":"1","key":"e_1_2_7_18_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3522618","article-title":"Real-time style modelling of human locomotion via feature-wise transformations and local motion phases","volume":"5","author":"Mason I.","year":"2022","journal-title":"Proceedings of the ACM on Computer Graphics and Interactive Techniques"},{"key":"e_1_2_7_19_2","doi-asserted-by":"crossref","unstructured":"MaedaT. UkitaN.: Motionaug: Augmentation with physical correction for human motion prediction. InProceedings of the IEEE\/CVF conference on computer vision and pattern recognition(2022) pp.6427\u20136436. 6","DOI":"10.1109\/CVPR52688.2022.00632"},{"key":"e_1_2_7_20_2","unstructured":"NaS. YooS. ChooJ.: Miso: Mutual information loss with stochastic style representations for multimodal image-to-image translation.arXiv preprint arXiv:1902.03938(2019). 4"},{"issue":"3","key":"e_1_2_7_21_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3480145","article-title":"Diverse motion stylization for multiple style domains via spatial-temporal graph-based generative model","volume":"4","author":"Park S.","year":"2021","journal-title":"Proceedings of the ACM on Computer Graphics and Interactive Techniques"},{"key":"e_1_2_7_22_2","doi-asserted-by":"crossref","DOI":"10.1609\/aaai.v32i1.11671","article-title":"Film: Visual reasoning with a general conditioning layer","volume":"32","author":"Perez E.","year":"2018","journal-title":"Proceedings of the AAAI conference on artificial intelligence"},{"key":"e_1_2_7_23_2","doi-asserted-by":"crossref","unstructured":"RaabS. GatI. SalaN. TevetG. Shalev-ArkushinR. FriedO. BermanoA. H. Cohen-OrD.: Monkey see monkey do: Harnessing self-attention in motion diffusion for zero-shot motion transfer. InSIGGRAPH Asia 2024 Conference Papers(2024) pp.1\u201313. 2","DOI":"10.1145\/3680528.3687579"},{"issue":"142","key":"e_1_2_7_24_2","first-page":"1","article-title":"Coding theorems for a discrete source with a fidelity criterion","volume":"4","author":"Shannon C. E.","year":"1959","journal-title":"IRE Nat. Conv. Rec"},{"key":"e_1_2_7_25_2","doi-asserted-by":"crossref","unstructured":"SongW. JinX. LiS. ChenC. HaoA. HouX. LiN. QinH.: Arbitrary motion style transfer with multi-condition motion latent diffusion model. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2024) pp.821\u2013830. 2","DOI":"10.1109\/CVPR52733.2024.00084"},{"issue":"4","key":"e_1_2_7_26_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3528223.3530178","article-title":"Deepphase: Periodic autoencoders for learning motion phase manifolds","volume":"41","author":"Starke S.","year":"2022","journal-title":"ACM Transactions on Graphics (ToG)"},{"key":"e_1_2_7_27_2","first-page":"48382","volume-title":"Advances in Neural Information Processing Systems","author":"Tian Y.","year":"2023"},{"key":"e_1_2_7_28_2","unstructured":"TevetG. RaabS. GordonB. ShafirY. Cohen-orD. BermanoA. H.: Human motion diffusion model. InThe Eleventh International Conference on Learning Representations(2023). 2"},{"key":"e_1_2_7_29_2","doi-asserted-by":"crossref","unstructured":"TangX. WuL. WangH. HuB. GongX. LiaoY. LiS. KouQ. JinX.: Rsmt: Real-time stylized motion transition for characters. InACM SIGGRAPH 2023 Conference Proceedings(2023) pp.1\u201310. 2","DOI":"10.1145\/3588432.3591514"},{"key":"e_1_2_7_30_2","doi-asserted-by":"crossref","unstructured":"TangX. WuL. WangH. WuY. HuB. LiS. GongX. LiaoY. KouQ. JinX.: Decoupling contact for fine-grained motion style transfer. InSIGGRAPH Asia 2024 Conference Papers(2024) pp.1\u201311. 2","DOI":"10.1145\/3680528.3687609"},{"key":"e_1_2_7_31_2","doi-asserted-by":"crossref","unstructured":"TaoT. ZhanX. ChenZ. van dePanneM.: Styleerd: Responsive and coherent online motion style transfer. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2022) pp.6593\u20136603. 10","DOI":"10.1109\/CVPR52688.2022.00648"},{"issue":"11","key":"e_1_2_7_32_2","article-title":"Visualizing data using t-sne","volume":"9","author":"Van der Maaten L.","year":"2008","journal-title":"Journal of machine learning research"},{"key":"e_1_2_7_33_2","article-title":"Neural discrete representation learning","volume":"30","author":"Van Den Oord A.","year":"2017","journal-title":"Advances in neural information processing systems"},{"issue":"4","key":"e_1_2_7_34_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2766999","article-title":"Realtime style transfer for unlabeled heterogeneous human motion","volume":"34","author":"Xia S.","year":"2015","journal-title":"ACM Transactions on Graphics (TOG)"},{"issue":"4","key":"e_1_2_7_35_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3658137","article-title":"Moconvq: Unified physics-based motion control via scalable discrete representations","volume":"43","author":"Yao H.","year":"2024","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_7_36_2","unstructured":"ZhangL. RaoA. AgrawalaM.: Adding conditional control to text-to-image diffusion models. InProceedings of the IEEE\/CVF international conference on computer vision(2023) pp.3836\u20133847. 2"},{"key":"e_1_2_7_37_2","first-page":"405","volume-title":"European Conference on Computer Vision","author":"Zhong L.","year":"2024"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70377","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70377","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70377","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T15:14:13Z","timestamp":1777562053000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70377"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,30]]},"references-count":36,"alternative-id":["10.1111\/cgf.70377"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70377","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,30]]},"assertion":[{"value":"2026-04-30","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70377"}}