{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T19:06:30Z","timestamp":1783105590777,"version":"3.54.6"},"reference-count":95,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2021,12,1]],"date-time":"2021-12-01T00:00:00Z","timestamp":1638316800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"name":"The Swedish Research Council a.k.a. Vetenskapsr\u00e5det","award":["grants no. 2018-05409 and 2019-03694"],"award-info":[{"award-number":["grants no. 2018-05409 and 2019-03694"]}]},{"name":"The Knut and Alice Wallenberg Foundation","award":["WASP"],"award-info":[{"award-number":["WASP"]}]},{"name":"The Marianne and Marcus Wallenberg Foundation","award":["MMW 2020.0102"],"award-info":[{"award-number":["MMW 2020.0102"]}]},{"DOI":"10.13039\/501100010190","name":"GENCI","doi-asserted-by":"crossref","award":["allocation 2020-[A0091011996]"],"award-info":[{"award-number":["allocation 2020-[A0091011996]"]}],"id":[{"id":"10.13039\/501100010190","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:p>Dance requires skillful composition of complex movements that follow rhythmic, tonal and timbral features of music. Formally, generating dance conditioned on a piece of music can be expressed as a problem of modelling a high-dimensional continuous motion signal, conditioned on an audio signal. In this work we make two contributions to tackle this problem. First, we present a novel probabilistic autoregressive architecture that models the distribution over future poses with a normalizing flow conditioned on previous poses as well as music context, using a multimodal transformer encoder. Second, we introduce the currently largest 3D dance-motion dataset, obtained with a variety of motion-capture technologies, and including both professional and casual dancers. Using this dataset, we compare our new model against two baselines, via objective metrics and a user study, and show that both the ability to model a probability distribution, as well as being able to attend over a large motion and music context are necessary to produce interesting, diverse, and realistic dance that matches the music.<\/jats:p>","DOI":"10.1145\/3478513.3480570","type":"journal-article","created":{"date-parts":[[2021,12,10]],"date-time":"2021-12-10T18:29:20Z","timestamp":1639160960000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":70,"title":["Transflower"],"prefix":"10.1145","volume":"40","author":[{"given":"Guillermo","family":"Valle-P\u00e9rez","sequence":"first","affiliation":[{"name":"University of Bordeaux, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gustav Eje","family":"Henter","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jonas","family":"Beskow","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andre","family":"Holzapfel","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pierre-Yves","family":"Oudeyer","sequence":"additional","affiliation":[{"name":"University of Bordeaux, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Simon","family":"Alexanderson","sequence":"additional","affiliation":[{"name":"KTH Royal Institute of Technology, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,12,10]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"Josh Abramson Arun Ahuja Iain Barr Arthur Brussee Federico Carnevale Mary Cassin Rachita Chhaparia Stephen Clark Bogdan Damoc Andrew Dudzik etal 2020. Imitating interactive intelligence. arXiv preprint arXiv:2012.05672 (2020).  Josh Abramson Arun Ahuja Iain Barr Arthur Brussee Federico Carnevale Mary Cassin Rachita Chhaparia Stephen Clark Bogdan Damoc Andrew Dudzik et al. 2020. Imitating interactive intelligence. arXiv preprint arXiv:2012.05672 (2020)."},{"key":"e_1_2_2_2_1","volume-title":"Groovenet: Real-time music-driven dance movement generation using artificial neural networks. networks 8, 17","author":"Alemi Omid","year":"2017","unstructured":"Omid Alemi , Jules Fran\u00e7oise , and Philippe Pasquier . 2017 . Groovenet: Real-time music-driven dance movement generation using artificial neural networks. networks 8, 17 (2017), 26. Omid Alemi, Jules Fran\u00e7oise, and Philippe Pasquier. 2017. Groovenet: Real-time music-driven dance movement generation using artificial neural networks. networks 8, 17 (2017), 26."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.13946"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/566654.566606"},{"key":"e_1_2_2_5_1","unstructured":"Jody Avirgan. 2013. Why Spiderman is Such a Good Dancer. https:\/\/www.wnycstudios.org\/podcasts\/radiolab\/articles\/299399-why-spiderman-such-good-dancer  Jody Avirgan. 2013. Why Spiderman is Such a Good Dancer. https:\/\/www.wnycstudios.org\/podcasts\/radiolab\/articles\/299399-why-spiderman-such-good-dancer"},{"key":"e_1_2_2_6_1","volume-title":"Neurocognitive control in dance perception and performance. Acta psychologica 139, 2","author":"Bl\u00e4sing Bettina","year":"2012","unstructured":"Bettina Bl\u00e4sing , Beatriz Calvo-Merino , Emily S Cross , Corinne Jola , Juliane Honisch , and Catherine J Stevens . 2012. Neurocognitive control in dance perception and performance. Acta psychologica 139, 2 ( 2012 ), 300--308. Bettina Bl\u00e4sing, Beatriz Calvo-Merino, Emily S Cross, Corinne Jola, Juliane Honisch, and Catherine J Stevens. 2012. Neurocognitive control in dance perception and performance. Acta psychologica 139, 2 (2012), 300--308."},{"key":"e_1_2_2_7_1","volume-title":"Proc. Int. Conf. Digital Audio Effects. 135--139","author":"B\u00f6ck Sebastian","year":"2011","unstructured":"Sebastian B\u00f6ck and Markus Schedl . 2011 . Enhanced beat tracking with context-aware neural networks . In Proc. Int. Conf. Digital Audio Effects. 135--139 . Sebastian B\u00f6ck and Markus Schedl. 2011. Enhanced beat tracking with context-aware neural networks. In Proc. Int. Conf. Digital Audio Effects. 135--139."},{"key":"e_1_2_2_8_1","volume-title":"Black","author":"Bogo Federica","year":"2016","unstructured":"Federica Bogo , Angjoo Kanazawa , Christoph Lassner , Peter Gehler , Javier Romero , and Michael J . Black . 2016 . Keep it SMPL : Automatic estimation of 3D human pose and shape from a single image. In European Conference on Computer Vision. Springer , 561--578. Federica Bogo, Angjoo Kanazawa, Christoph Lassner, Peter Gehler, Javier Romero, and Michael J. Black. 2016. Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image. In European Conference on Computer Vision. Springer, 561--578."},{"key":"e_1_2_2_9_1","volume-title":"GANs, Normalizing Flows, Energy-Based and Autoregressive Models. arXiv preprint arXiv:2103.04922","author":"Bond-Taylor Sam","year":"2021","unstructured":"Sam Bond-Taylor , Adam Leach , Yang Long , and Chris G Willcocks . 2021. Deep Generative Modelling: A Comparative Review of VAEs , GANs, Normalizing Flows, Energy-Based and Autoregressive Models. arXiv preprint arXiv:2103.04922 ( 2021 ). Sam Bond-Taylor, Adam Leach, Yang Long, and Chris G Willcocks. 2021. Deep Generative Modelling: A Comparative Review of VAEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models. arXiv preprint arXiv:2103.04922 (2021)."},{"key":"e_1_2_2_10_1","volume-title":"Large Scale GAN Training for High Fidelity Natural Image Synthesis. In International Conference on Learning Representations.","author":"Brock Andrew","year":"2018","unstructured":"Andrew Brock , Jeff Donahue , and Karen Simonyan . 2018 . Large Scale GAN Training for High Fidelity Natural Image Synthesis. In International Conference on Learning Representations. Andrew Brock, Jeff Donahue, and Karen Simonyan. 2018. Large Scale GAN Training for High Fidelity Natural Image Synthesis. In International Conference on Learning Representations."},{"key":"e_1_2_2_11_1","first-page":"1877","article-title":"Language Models are Few-Shot Learners","volume":"33","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown , Benjamin Mann , Nick Ryder , Melanie Subbiah , Jared D. Kaplan , Prafulla Dhariwal , Arvind Neelakantan , Pranav Shyam , Girish Sastry , Amanda Askell , Sandhini Agarwal , Ariel Herbert-Voss , Gretchen Krueger , Tom Henighan , Rewon Child , Aditya Ramesh , Daniel Ziegler , Jeffrey Wu , Clemens Winter , Chris Hesse , Mark Chen , Eric Sigler , Mateusz Litwin , Scott Gray , Benjamin Chess , Jack Clark , Christopher Berner , Sam McCandlish , Alec Radford , Ilya Sutskever , and Dario Amodei . 2020 . Language Models are Few-Shot Learners . In Proc. NeurIPS , Vol. 33. 1877 -- 1901 . https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Proc. NeurIPS, Vol. 33. 1877--1901. https:\/\/proceedings.neurips.cc\/paper\/2020\/file\/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf","journal-title":"Proc. NeurIPS"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.3389\/fpsyg.2013.00183"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.173"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/977399.977857"},{"key":"e_1_2_2_15_1","unstructured":"CMU Graphics Lab. 2003. Carnegie Mellon University motion capture database. http:\/\/mocap.cs.cmu.edu\/  CMU Graphics Lab. 2003. Carnegie Mellon University motion capture database. http:\/\/mocap.cs.cmu.edu\/"},{"key":"e_1_2_2_16_1","volume-title":"Generative choreography using deep learning. arXiv preprint arXiv:1605.06921","author":"Crnkovic-Friis Luka","year":"2016","unstructured":"Luka Crnkovic-Friis and Louise Crnkovic-Friis . 2016. Generative choreography using deep learning. arXiv preprint arXiv:1605.06921 ( 2016 ). Luka Crnkovic-Friis and Louise Crnkovic-Friis. 2016. Generative choreography using deep learning. arXiv preprint arXiv:1605.06921 (2016)."},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201371"},{"key":"e_1_2_2_18_1","volume-title":"Alec Radford, and Ilya Sutskever.","author":"Dhariwal Prafulla","year":"2020","unstructured":"Prafulla Dhariwal , Heewoo Jun , Christine Payne , Jong Wook Kim , Alec Radford, and Ilya Sutskever. 2020 . Jukebox : A generative model for music. arXiv preprint arXiv:2005.00341 (2020). Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. 2020. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341 (2020)."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305381.3305489"},{"key":"e_1_2_2_20_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly etal 2020. An image is worth 16\u00d716 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020).  Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16\u00d716 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)."},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2011.73"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3267851.3267898"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/2919332.2919834"},{"key":"e_1_2_2_24_1","volume-title":"Proceedings of SMC","author":"Fukayama Satoru","year":"2015","unstructured":"Satoru Fukayama and Masataka Goto . 2015 . Music content driven automated choreography with beat-wise motion connectivity constraints . Proceedings of SMC (2015), 177--183. Satoru Fukayama and Masataka Goto. 2015. Music content driven automated choreography with beat-wise motion connectivity constraints. Proceedings of SMC (2015), 177--183."},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1080\/10867651.1998.10487493"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/1186562.1015755"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.5244\/C.31.119"},{"key":"e_1_2_2_28_1","volume-title":"Learning Speech-driven 3D Conversational Gestures from Video. arXiv preprint arXiv:2102.06837","author":"Habibie Ikhsanul","year":"2021","unstructured":"Ikhsanul Habibie , Weipeng Xu , Dushyant Mehta , Lingjie Liu , Hans-Peter Seidel , Gerard Pons-Moll , Mohamed Elgharib , and Christian Theobalt . 2021. Learning Speech-driven 3D Conversational Gestures from Video. arXiv preprint arXiv:2102.06837 ( 2021 ). Ikhsanul Habibie, Weipeng Xu, Dushyant Mehta, Lingjie Liu, Hans-Peter Seidel, Gerard Pons-Moll, Mohamed Elgharib, and Christian Theobalt. 2021. Learning Speech-driven 3D Conversational Gestures from Video. arXiv preprint arXiv:2102.06837 (2021)."},{"key":"e_1_2_2_29_1","volume-title":"Studying rhythmical structures in Norwegian folk music and dance using motion capture technology: A case study of Norwegian telespringar. Musikk og Tradisjon 28","author":"Haugen Mari Romarheim","year":"2014","unstructured":"Mari Romarheim Haugen . 2014. Studying rhythmical structures in Norwegian folk music and dance using motion capture technology: A case study of Norwegian telespringar. Musikk og Tradisjon 28 ( 2014 ), 27--52. Mari Romarheim Haugen. 2014. Studying rhythmical structures in Norwegian folk music and dance using motion capture technology: A case study of Norwegian telespringar. Musikk og Tradisjon 28 (2014), 27--52."},{"key":"e_1_2_2_30_1","unstructured":"Tom Henighan Jared Kaplan Mor Katz Mark Chen Christopher Hesse Jacob Jackson Heewoo Jun Tom B Brown Prafulla Dhariwal Scott Gray etal 2020. Scaling laws for autoregressive generative modeling. arXiv preprint arXiv:2010.14701 (2020).  Tom Henighan Jared Kaplan Mor Katz Mark Chen Christopher Hesse Jacob Jackson Heewoo Jun Tom B Brown Prafulla Dhariwal Scott Gray et al. 2020. Scaling laws for autoregressive generative modeling. arXiv preprint arXiv:2010.14701 (2020)."},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417836"},{"key":"e_1_2_2_32_1","volume-title":"International Conference on Machine Learning. PMLR, 2722--2730","author":"Ho Jonathan","year":"2019","unstructured":"Jonathan Ho , Xi Chen , Aravind Srinivas , Yan Duan , and Pieter Abbeel . 2019 . Flow++: Improving flow-based generative models with variational dequantization and architecture design . In International Conference on Machine Learning. PMLR, 2722--2730 . Jonathan Ho, Xi Chen, Aravind Srinivas, Yan Duan, and Pieter Abbeel. 2019. Flow++: Improving flow-based generative models with variational dequantization and architecture design. In International Conference on Machine Learning. PMLR, 2722--2730."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392440"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073663"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925975"},{"key":"e_1_2_2_36_1","volume-title":"The Curious Case of Neural Text Degeneration. In International Conference on Learning Representations.","author":"Holtzman Ari","year":"2020","unstructured":"Ari Holtzman , Jan Buys , Li Du , Maxwell Forbes , and Yejin Choi . 2020 . The Curious Case of Neural Text Degeneration. In International Conference on Learning Representations. Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The Curious Case of Neural Text Degeneration. In International Conference on Learning Representations."},{"key":"e_1_2_2_37_1","volume-title":"Proceedings of the ICTM Study Group on Sound, Movement, and the Sciences Symposium.","author":"Holzapfel Andre","year":"2020","unstructured":"Andre Holzapfel , Michael Hagleitner , and Stella Pashalidou . 2020 . Diversity of Traditional Dance Expression in Crete: Data Collection, Research Questions, and Method Development . In Proceedings of the ICTM Study Group on Sound, Movement, and the Sciences Symposium. Andre Holzapfel, Michael Hagleitner, and Stella Pashalidou. 2020. Diversity of Traditional Dance Expression in Crete: Data Collection, Research Questions, and Method Development. In Proceedings of the ICTM Study Group on Sound, Movement, and the Sciences Symposium."},{"key":"e_1_2_2_38_1","volume-title":"Music Transformer: Generating Music with Long-Term Structure. In International Conference on Learning Representations.","author":"Anna Huang Cheng-Zhi","year":"2018","unstructured":"Cheng-Zhi Anna Huang , Ashish Vaswani , Jakob Uszkoreit , Ian Simon , Curtis Hawthorne , Noam Shazeer , Andrew M Dai , Matthew D Hoffman , Monica Dinculescu , and Douglas Eck . 2018 . Music Transformer: Generating Music with Long-Term Structure. In International Conference on Learning Representations. Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck. 2018. Music Transformer: Generating Music with Long-Term Structure. In International Conference on Learning Representations."},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459932"},{"key":"e_1_2_2_40_1","volume-title":"Scaling laws for neural language models. arXiv preprint arXiv:2001.08361","author":"Kaplan Jared","year":"2020","unstructured":"Jared Kaplan , Sam McCandlish , Tom Henighan , Tom B. Brown , Benjamin Chess , Rewon Child , Scott Gray , Alec Radford , Jeffrey Wu , and Dario Amodei . 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361 ( 2020 ). Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361 (2020)."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00453"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.5555\/3327546.3327685"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/1186562.1015760"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/566654.566605"},{"key":"e_1_2_2_45_1","unstructured":"Florian Krebs Sebastian B\u00f6ck and Gerhard Widmer. 2015. An Efficient State-Space Model for Joint Tempo and Meter Tracking.. In ISMIR. 72--78.  Florian Krebs Sebastian B\u00f6ck and Gerhard Widmer. 2015. An Efficient State-Space Model for Joint Tempo and Meter Tracking.. In ISMIR. 72--78."},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397481.3450692"},{"key":"e_1_2_2_47_1","unstructured":"OxAI Labs. 2019. DeepSaber. https:\/\/github.com\/oxai\/deepsaber\/.  OxAI Labs. 2019. DeepSaber. https:\/\/github.com\/oxai\/deepsaber\/."},{"key":"e_1_2_2_48_1","volume-title":"The dancing species: how moving together in time helps make us human. Aeon (June","author":"LaMothe Kimerer","year":"2019","unstructured":"Kimerer LaMothe . 2019. The dancing species: how moving together in time helps make us human. Aeon (June 2019 ). https:\/\/aeon.co\/ideas\/the-dancing-species-how-moving-together-in-time-helps-make-us-human Kimerer LaMothe. 2019. The dancing species: how moving together in time helps make us human. Aeon (June 2019). https:\/\/aeon.co\/ideas\/the-dancing-species-how-moving-together-in-time-helps-make-us-human"},{"key":"e_1_2_2_49_1","unstructured":"Ben Lang. 2021. The Future is Now: Live Breakdance Battles in VR Are Connecting People Across the Globe. https:\/\/www.roadtovr.com\/vr-dance-battle-vrchat-breakdance\/  Ben Lang. 2021. The Future is Now: Live Breakdance Battles in VR Are Connecting People Across the Globe. https:\/\/www.roadtovr.com\/vr-dance-battle-vrchat-breakdance\/"},{"key":"e_1_2_2_50_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision. 763--772","author":"Lee Gilwoo","year":"2019","unstructured":"Gilwoo Lee , Zhiwei Deng , Shugao Ma , Takaaki Shiratori , Siddhartha S Srinivasa , and Yaser Sheikh . 2019 a. Talking with hands 16.2 m: A large-scale dataset of synchronized body-finger motion and audio for conversational motion analysis and synthesis . In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 763--772 . Gilwoo Lee, Zhiwei Deng, Shugao Ma, Takaaki Shiratori, Siddhartha S Srinivasa, and Yaser Sheikh. 2019a. Talking with hands 16.2 m: A large-scale dataset of synchronized body-finger motion and audio for conversational motion analysis and synthesis. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 763--772."},{"key":"e_1_2_2_51_1","volume-title":"Dancing to music. arXiv preprint arXiv:1911.02001","author":"Lee Hsin-Ying","year":"2019","unstructured":"Hsin-Ying Lee , Xiaodong Yang , Ming-Yu Liu , Ting-Chun Wang , Yu-Ding Lu , Ming-Hsuan Yang , and Jan Kautz . 2019b. Dancing to music. arXiv preprint arXiv:1911.02001 ( 2019 ). Hsin-Ying Lee, Xiaodong Yang, Ming-Yu Liu, Ting-Chun Wang, Yu-Ding Lu, Ming-Hsuan Yang, and Jan Kautz. 2019b. Dancing to music. arXiv preprint arXiv:1911.02001 (2019)."},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/566654.566607"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185524"},{"key":"e_1_2_2_54_1","volume-title":"DanceNet3D: Music Based Dance Generation with Parametric Motion Transformer. arXiv preprint arXiv:2103.10206","author":"Li Buyu","year":"2021","unstructured":"Buyu Li , Yongchi Zhao , and Lu Sheng . 2021b. DanceNet3D: Music Based Dance Generation with Parametric Motion Transformer. arXiv preprint arXiv:2103.10206 ( 2021 ). Buyu Li, Yongchi Zhao, and Lu Sheng. 2021b. DanceNet3D: Music Based Dance Generation with Parametric Motion Transformer. arXiv preprint arXiv:2103.10206 (2021)."},{"key":"e_1_2_2_55_1","volume-title":"Learning to Generate Diverse Dance Motions with Transformer. arXiv preprint arXiv:2008.08171","author":"Li Jiaman","year":"2020","unstructured":"Jiaman Li , Yihang Yin , Hang Chu , Yi Zhou , Tingwu Wang , Sanja Fidler , and Hao Li. 2020. Learning to Generate Diverse Dance Motions with Transformer. arXiv preprint arXiv:2008.08171 ( 2020 ). Jiaman Li, Yihang Yin, Hang Chu, Yi Zhou, Tingwu Wang, Sanja Fidler, and Hao Li. 2020. Learning to Generate Diverse Dance Motions with Transformer. arXiv preprint arXiv:2008.08171 (2020)."},{"key":"e_1_2_2_56_1","volume-title":"Learn to Dance with AIST++: Music Conditioned 3D Dance Generation. arXiv preprint arXiv:2101.08779","author":"Li Ruilong","year":"2021","unstructured":"Ruilong Li , Shan Yang , David A Ross , and Angjoo Kanazawa . 2021a. Learn to Dance with AIST++: Music Conditioned 3D Dance Generation. arXiv preprint arXiv:2101.08779 ( 2021 ). Ruilong Li, Shan Yang, David A Ross, and Angjoo Kanazawa. 2021a. Learn to Dance with AIST++: Music Conditioned 3D Dance Generation. arXiv preprint arXiv:2101.08779 (2021)."},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392422"},{"key":"e_1_2_2_58_1","unstructured":"lox9973. 2021. ShaderMotion. https:\/\/gitlab.com\/lox9973\/ShaderMotion.  lox9973. 2021. ShaderMotion. https:\/\/gitlab.com\/lox9973\/ShaderMotion."},{"key":"e_1_2_2_59_1","volume-title":"International Conference on Computer Vision. 5442--5451","author":"Mahmood Naureen","unstructured":"Naureen Mahmood , Nima Ghorbani , Nikolaus F. Troje , Gerard Pons-Moll , and Michael J. Black . 2019. AMASS: Archive of Motion Capture as Surface Shapes . In International Conference on Computer Vision. 5442--5451 . Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. 2019. AMASS: Archive of Motion Capture as Surface Shapes. In International Conference on Computer Vision. 5442--5451."},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICAR.2015.7251476"},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neuron.2020.09.017"},{"key":"e_1_2_2_62_1","volume-title":"Learning human behaviors from motion capture by adversarial imitation. arXiv preprint arXiv:1707.02201","author":"Merel Josh","year":"2017","unstructured":"Josh Merel , Yuval Tassa , Sriram Srinivasan , Jay Lemmon , Ziyu Wang , Greg Wayne , and Nicolas Heess . 2017. Learning human behaviors from motion capture by adversarial imitation. arXiv preprint arXiv:1707.02201 ( 2017 ). Josh Merel, Yuval Tassa, Sriram Srinivasan, Jay Lemmon, Ziyu Wang, Greg Wayne, and Nicolas Heess. 2017. Learning human behaviors from motion capture by adversarial imitation. arXiv preprint arXiv:1707.02201 (2017)."},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1177\/0956797612446707"},{"key":"e_1_2_2_64_1","volume-title":"Sound & Music Computing Conference.","author":"Misgeld Olof","year":"2019","unstructured":"Olof Misgeld , Andre Holzapfel , and Sven Ahlb\u00e4ck . 2019 . Dancing Dots - Investigating the Link between Dancer and Musician in Swedish Folk Dance . In Sound & Music Computing Conference. Olof Misgeld, Andre Holzapfel, and Sven Ahlb\u00e4ck. 2019. Dancing Dots - Investigating the Link between Dancer and Musician in Swedish Folk Dance. In Sound & Music Computing Conference."},{"key":"e_1_2_2_65_1","volume-title":"Hypotheses on the choreographic roots of the musical meter: a case study on Afro-Brazilian dance and music. Debates actuales en evoluci\u00f3n, desarrollo y cognici\u00f3n e implicancias socio-culturales","author":"Naveda Luiz","year":"2011","unstructured":"Luiz Naveda and Marc Leman . 2011. Hypotheses on the choreographic roots of the musical meter: a case study on Afro-Brazilian dance and music. Debates actuales en evoluci\u00f3n, desarrollo y cognici\u00f3n e implicancias socio-culturales ( 2011 ), 477--495. Luiz Naveda and Marc Leman. 2011. Hypotheses on the choreographic roots of the musical meter: a case study on Afro-Brazilian dance and music. Debates actuales en evoluci\u00f3n, desarrollo y cognici\u00f3n e implicancias socio-culturales (2011), 477--495."},{"key":"e_1_2_2_66_1","first-page":"1","article-title":"Normalizing Flows for Probabilistic Modeling and Inference","volume":"22","author":"Papamakarios George","year":"2021","unstructured":"George Papamakarios , Eric Nalisnick , Danilo Jimenez Rezende , Shakir Mohamed , and Balaji Lakshminarayanan . 2021 . Normalizing Flows for Probabilistic Modeling and Inference . Journal of Machine Learning Research 22 , 57 (2021), 1 -- 64 . http:\/\/jmlr.org\/papers\/v22\/19-1028.html George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, and Balaji Lakshminarayanan. 2021. Normalizing Flows for Probabilistic Modeling and Inference. Journal of Machine Learning Research 22, 57 (2021), 1--64. http:\/\/jmlr.org\/papers\/v22\/19-1028.html","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_2_67_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00244"},{"key":"e_1_2_2_68_1","volume-title":"Proceedings of the British Machine Vision Conference (BMVC'18)","author":"Pavllo Dario","year":"2018","unstructured":"Dario Pavllo , David Grangier , and Michael Auli . 2018 . QuaterNet: A quaternion-based recurrent model for human motion . In Proceedings of the British Machine Vision Conference (BMVC'18) . BMVA Press, Durham, UK, 14 pages. http:\/\/www.bmva.org\/bmvc\/ 2018\/contents\/papers\/0675.pdf Dario Pavllo, David Grangier, and Michael Auli. 2018. QuaterNet: A quaternion-based recurrent model for human motion. In Proceedings of the British Machine Vision Conference (BMVC'18). BMVA Press, Durham, UK, 14 pages. http:\/\/www.bmva.org\/bmvc\/2018\/contents\/papers\/0675.pdf"},{"key":"e_1_2_2_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201311"},{"key":"e_1_2_2_70_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3272127.3275014","article-title":"Sfv: Reinforcement learning of physical skills from videos","volume":"37","author":"Peng Xue Bin","year":"2018","unstructured":"Xue Bin Peng , Angjoo Kanazawa , Jitendra Malik , Pieter Abbeel , and Sergey Levine . 2018 b. Sfv: Reinforcement learning of physical skills from videos . ACM Transactions On Graphics (TOG) 37 , 6 (2018), 1 -- 14 . Xue Bin Peng, Angjoo Kanazawa, Jitendra Malik, Pieter Abbeel, and Sergey Levine. 2018b. Sfv: Reinforcement learning of physical skills from videos. ACM Transactions On Graphics (TOG) 37, 6 (2018), 1--14.","journal-title":"ACM Transactions On Graphics (TOG)"},{"key":"e_1_2_2_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/3450626.3459670"},{"key":"e_1_2_2_72_1","volume-title":"Action-Conditioned 3D Human Motion Synthesis with Transformer VAE. arXiv preprint arXiv:2104.05670","author":"Petrovich Mathis","year":"2021","unstructured":"Mathis Petrovich , Michael J Black , and G\u00fcl Varol . 2021. Action-Conditioned 3D Human Motion Synthesis with Transformer VAE. arXiv preprint arXiv:2104.05670 ( 2021 ). Mathis Petrovich, Michael J Black, and G\u00fcl Varol. 2021. Action-Conditioned 3D Human Motion Synthesis with Transformer VAE. arXiv preprint arXiv:2104.05670 (2021)."},{"key":"e_1_2_2_73_1","doi-asserted-by":"publisher","DOI":"10.1098\/rstb.2020.0334"},{"key":"e_1_2_2_74_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683143"},{"key":"e_1_2_2_75_1","volume-title":"Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683","author":"Raffel Colin","year":"2019","unstructured":"Colin Raffel , Noam Shazeer , Adam Roberts , Katherine Lee , Sharan Narang , Michael Matena , Yanqi Zhou , Wei Li , and Peter J Liu . 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683 ( 2019 ). Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683 (2019)."},{"key":"e_1_2_2_76_1","volume-title":"Zero-shot text-to-image generation. arXiv preprint arXiv:2102.12092","author":"Ramesh Aditya","year":"2021","unstructured":"Aditya Ramesh , Mikhail Pavlov , Gabriel Goh , Scott Gray , Chelsea Voss , Alec Radford , Mark Chen , and Ilya Sutskever . 2021. Zero-shot text-to-image generation. arXiv preprint arXiv:2102.12092 ( 2021 ). Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. arXiv preprint arXiv:2102.12092 (2021)."},{"key":"e_1_2_2_77_1","volume-title":"Prompt programming for large language models: Beyond the few-shot paradigm. arXiv preprint arXiv:2102.07350","author":"Reynolds Laria","year":"2021","unstructured":"Laria Reynolds and Kyle McDonell . 2021. Prompt programming for large language models: Beyond the few-shot paradigm. arXiv preprint arXiv:2102.07350 ( 2021 ). Laria Reynolds and Kyle McDonell. 2021. Prompt programming for large language models: Beyond the few-shot paradigm. arXiv preprint arXiv:2102.07350 (2021)."},{"key":"e_1_2_2_78_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW54120.2021.00201"},{"key":"e_1_2_2_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/1276377.1276510"},{"key":"e_1_2_2_80_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461368"},{"key":"e_1_2_2_81_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386569.3392450"},{"key":"e_1_2_2_82_1","volume-title":"Augmented reality (AR) and virtual reality (VR) headset shipments worldwide from 2020 to","year":"2025","unstructured":"Statista. 2020. Augmented reality (AR) and virtual reality (VR) headset shipments worldwide from 2020 to 2025 . https:\/\/www.statista.com\/statistics\/653390\/worldwide-virtual-and-augmented-reality-headset-shipments\/ Statista. 2020. Augmented reality (AR) and virtual reality (VR) headset shipments worldwide from 2020 to 2025. https:\/\/www.statista.com\/statistics\/653390\/worldwide-virtual-and-augmented-reality-headset-shipments\/"},{"key":"e_1_2_2_83_1","doi-asserted-by":"publisher","DOI":"10.7210\/jrsj.28.723"},{"key":"e_1_2_2_84_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240526"},{"key":"e_1_2_2_85_1","doi-asserted-by":"publisher","DOI":"10.1525\/mp.2010.28.1.59"},{"key":"e_1_2_2_86_1","doi-asserted-by":"publisher","DOI":"10.1167\/2.5.2"},{"key":"e_1_2_2_87_1","volume-title":"Proceedings of the 20th International Society for Music Information Retrieval Conference, ISMIR 2019","author":"Tsuchida Shuhei","year":"2019","unstructured":"Shuhei Tsuchida , Satoru Fukayama , Masahiro Hamasaki , and Masataka Goto . 2019 . AIST Dance Video Database: Multi-genre, Multi-dancer, and Multi-camera Database for Dance Information Processing . In Proceedings of the 20th International Society for Music Information Retrieval Conference, ISMIR 2019 . Delft, Netherlands, 501--510. Shuhei Tsuchida, Satoru Fukayama, Masahiro Hamasaki, and Masataka Goto. 2019. AIST Dance Video Database: Multi-genre, Multi-dancer, and Multi-camera Database for Dance Information Processing. In Proceedings of the 20th International Society for Music Information Retrieval Conference, ISMIR 2019. Delft, Netherlands, 501--510."},{"key":"e_1_2_2_88_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_2_2_89_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2007.1167"},{"key":"e_1_2_2_90_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-short.18"},{"key":"e_1_2_2_91_1","volume-title":"GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions. arXiv preprint arXiv:2104.14806","author":"Wu Chenfei","year":"2021","unstructured":"Chenfei Wu , Lun Huang , Qianxi Zhang , Binyang Li , Lei Ji , Fan Yang , Guillermo Sapiro , and Nan Duan . 2021 . GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions. arXiv preprint arXiv:2104.14806 (2021). Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan. 2021. GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions. arXiv preprint arXiv:2104.14806 (2021)."},{"key":"e_1_2_2_92_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3414005"},{"key":"e_1_2_2_93_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8793720"},{"key":"e_1_2_2_94_1","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR'18)","author":"Zhou Yi","year":"2018","unstructured":"Yi Zhou , Zimo Li , Shuangjiu Xiao , Chong He , Zeng Huang , and Hao Li . 2018 . Auto-conditioned recurrent networks for extended complex human motion synthesis . In Proceedings of the International Conference on Learning Representations (ICLR'18) . 13 pages. https:\/\/openreview.net\/forum?id=r11Q2SlRW Yi Zhou, Zimo Li, Shuangjiu Xiao, Chong He, Zeng Huang, and Hao Li. 2018. Auto-conditioned recurrent networks for extended complex human motion synthesis. In Proceedings of the International Conference on Learning Representations (ICLR'18). 13 pages. https:\/\/openreview.net\/forum?id=r11Q2SlRW"},{"key":"e_1_2_2_95_1","volume-title":"Music2Dance: DanceNet for Music-driven Dance Generation. arXiv e-prints","author":"Zhuang Wenlin","year":"2020","unstructured":"Wenlin Zhuang , Congyi Wang , Siyu Xia , Jinxiang Chai , and Yangang Wang . 2020. Music2Dance: DanceNet for Music-driven Dance Generation. arXiv e-prints ( 2020 ), arXiv-2002. Wenlin Zhuang, Congyi Wang, Siyu Xia, Jinxiang Chai, and Yangang Wang. 2020. Music2Dance: DanceNet for Music-driven Dance Generation. arXiv e-prints (2020), arXiv-2002."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3478513.3480570","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3478513.3480570","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:11:40Z","timestamp":1750191100000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3478513.3480570"}},"subtitle":["probabilistic autoregressive dance generation with multimodal attention"],"short-title":[],"issued":{"date-parts":[[2021,12]]},"references-count":95,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["10.1145\/3478513.3480570"],"URL":"https:\/\/doi.org\/10.1145\/3478513.3480570","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12]]},"assertion":[{"value":"2021-12-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}