{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T18:25:24Z","timestamp":1784399124703,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":52,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100003453","name":"Natural Science Foundation of Guangdong Province","doi-asserted-by":"publisher","award":["2019A1515010860, 2021A1515012301"],"award-info":[{"award-number":["2019A1515010860, 2021A1515012301"]}],"id":[{"id":"10.13039\/501100003453","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Natural Science Foundation of China","award":["62072191, 61972160"],"award-info":[{"award-number":["62072191, 61972160"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3547956","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:42:46Z","timestamp":1665416566000},"page":"5162-5171","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":51,"title":["Diverse Human Motion Prediction via Gumbel-Softmax Sampling from an Auxiliary Space"],"prefix":"10.1145","author":[{"given":"Lingwei","family":"Dang","sequence":"first","affiliation":[{"name":"South China University of Technology, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yongwei","family":"Nie","sequence":"additional","affiliation":[{"name":"South China University of Technology, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chengjiang","family":"Long","sequence":"additional","affiliation":[{"name":"Meta Reality Lab, Burlingame, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qing","family":"Zhang","sequence":"additional","affiliation":[{"name":"Sun Yat-sen University, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guiqing","family":"Li","sequence":"additional","affiliation":[{"name":"South China University of Technology, Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"please: A spatio-temporal transformer for 3d human motion prediction. arXiv preprint arXiv:2004.08692","author":"Aksan Emre","year":"2020","unstructured":"Emre Aksan , Peng Cao , Manuel Kaufmann , and Otmar Hilliges . 2020. Attention , please: A spatio-temporal transformer for 3d human motion prediction. arXiv preprint arXiv:2004.08692 , Vol. 2 , 3 ( 2020 ), 5. Emre Aksan, Peng Cao, Manuel Kaufmann, and Otmar Hilliges. 2020. Attention, please: A spatio-temporal transformer for 3d human motion prediction. arXiv preprint arXiv:2004.08692, Vol. 2, 3 (2020), 5."},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/3DV53792.2021.00066"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00724"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01114"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00527"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2018.00191"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00885"},{"key":"e_1_3_2_2_8_1","volume-title":"HiT-DVAE: Human Motion Generation via Hierarchical Transformer Dynamical VAE. arXiv preprint arXiv:2204.01565","author":"Bie Xiaoyu","year":"2022","unstructured":"Xiaoyu Bie , Wen Guo , Simon Leglaive , Lauren Girin , Francesc Moreno-Noguer , and Xavier Alameda-Pineda . 2022. HiT-DVAE: Human Motion Generation via Hierarchical Transformer Dynamical VAE. arXiv preprint arXiv:2204.01565 ( 2022 ). Xiaoyu Bie, Wen Guo, Simon Leglaive, Lauren Girin, Francesc Moreno-Noguer, and Xavier Alameda-Pineda. 2022. HiT-DVAE: Human Motion Generation via Hierarchical Transformer Dynamical VAE. arXiv preprint arXiv:2204.01565 (2022)."},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58571-6_14"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01144"},{"key":"e_1_3_2_2_11_1","volume-title":"2019 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 1423--1432","author":"Adeli Ehsan","year":"2019","unstructured":"Hsu-kuang Chiu, Ehsan Adeli , Borui Wang , De-An Huang , and Juan Carlos Niebles . 2019 . Action-agnostic human pose forecasting . In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 1423--1432 . Hsu-kuang Chiu, Ehsan Adeli, Borui Wang, De-An Huang, and Juan Carlos Niebles. 2019. Action-agnostic human pose forecasting. In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 1423--1432."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00702"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00477"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00655"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01127"},{"key":"e_1_3_2_2_16_1","volume-title":"Marta Garnelo, Matthew CH Lee, Hugh Salimbeni, Kai Arulkumaran, and Murray Shanahan.","author":"Dilokthanakul Nat","year":"2016","unstructured":"Nat Dilokthanakul , Pedro AM Mediano , Marta Garnelo, Matthew CH Lee, Hugh Salimbeni, Kai Arulkumaran, and Murray Shanahan. 2016 . Deep unsupervised clustering with gaussian mixture variational autoencoders. arXiv preprint arXiv:1611.02648 (2016). Nat Dilokthanakul, Pedro AM Mediano, Marta Garnelo, Matthew CH Lee, Hugh Salimbeni, Kai Arulkumaran, and Murray Shanahan. 2016. Deep unsupervised clustering with gaussian mixture variational autoencoders. arXiv preprint arXiv:1611.02648 (2016)."},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475439"},{"key":"e_1_3_2_2_18_1","volume-title":"Complementary Attention Gated Network for Pedestrian Trajectory Prediction. In AAAI Conference on Artificial Intelligence.","author":"Duan Jinghai","year":"2022","unstructured":"Jinghai Duan , Le Wang , Chengjiang Long , Sanping Zhou , Fang Zheng , Liushuai Shi , and Gang Hua . 2022 . Complementary Attention Gated Network for Pedestrian Trajectory Prediction. In AAAI Conference on Artificial Intelligence. Jinghai Duan, Le Wang, Chengjiang Long, Sanping Zhou, Fang Zheng, Liushuai Shi, and Gang Hua. 2022. Complementary Attention Gated Network for Pedestrian Trajectory Prediction. In AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/2919332.2919834"},{"key":"e_1_3_2_2_20_1","volume-title":"Generative adversarial nets. Advances in neural information processing systems","author":"Goodfellow Ian","year":"2014","unstructured":"Ian Goodfellow , Jean Pouget-Abadie , Mehdi Mirza , Bing Xu , David Warde-Farley , Sherjil Ozair , Aaron Courville , and Yoshua Bengio . 2014. Generative adversarial nets. Advances in neural information processing systems , Vol. 27 ( 2014 ). Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. Advances in neural information processing systems, Vol. 27 (2014)."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.525"},{"key":"e_1_3_2_2_22_1","volume-title":"6m: Large scale datasets and predictive methods for 3d human sensing in natural environments","author":"Ionescu Catalin","year":"2013","unstructured":"Catalin Ionescu , Dragos Papava , Vlad Olaru , and Cristian Sminchisescu . 2013. Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments . IEEE transactions on pattern analysis and machine intelligence, Vol. 36 , 7 ( 2013 ), 1325--1339. Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. 2013. Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE transactions on pattern analysis and machine intelligence, Vol. 36, 7 (2013), 1325--1339."},{"key":"e_1_3_2_2_23_1","volume-title":"Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114","author":"Kingma Diederik P","year":"2013","unstructured":"Diederik P Kingma and Max Welling . 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 ( 2013 ). Diederik P Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013)."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018553"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3038362"},{"key":"e_1_3_2_2_26_1","volume-title":"Symbiotic graph neural networks for 3d skeleton-based human action recognition and motion prediction","author":"Li Maosen","year":"2021","unstructured":"Maosen Li , Siheng Chen , Xu Chen , Ya Zhang , Yanfeng Wang , and Qi Tian . 2021. Symbiotic graph neural networks for 3d skeleton-based human action recognition and motion prediction . IEEE Transactions on Pattern Analysis and Machine Intelligence ( 2021 ). Maosen Li, Siheng Chen, Xu Chen, Ya Zhang, Yanfeng Wang, and Qi Tian. 2021. Symbiotic graph neural networks for 3d skeleton-based human action recognition and motion prediction. IEEE Transactions on Pattern Analysis and Machine Intelligence (2021)."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00029"},{"key":"e_1_3_2_2_28_1","volume-title":"Auto-conditioned recurrent networks for extended complex human motion synthesis. arXiv preprint arXiv:1707.05363","author":"Li Zimo","year":"2017","unstructured":"Zimo Li , Yi Zhou , Shuangjiu Xiao , Chong He , Zeng Huang , and Hao Li. 2017. Auto-conditioned recurrent networks for extended complex human motion synthesis. arXiv preprint arXiv:1707.05363 ( 2017 ). Zimo Li, Yi Zhou, Shuangjiu Xiao, Chong He, Zeng Huang, and Hao Li. 2017. Auto-conditioned recurrent networks for extended complex human motion synthesis. arXiv preprint arXiv:1707.05363 (2017)."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i3.16321"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01333"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01305"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3139918"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475630"},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00633"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58568-6_28"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01306"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00958"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01483-7"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.497"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW54120.2021.00257"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-019-08269-7"},{"key":"e_1_3_2_2_42_1","volume-title":"Social Interpretable Tree for Pedestrian Trajectory Prediction. In AAAI Conference on Artificial Intelligence.","author":"Shi Liushuai","year":"2022","unstructured":"Liushuai Shi , Le Wang , Chengjiang Long , Sanping Zhou , Fang Zheng , Nanning Zheng , and Gang Hua . 2022 . Social Interpretable Tree for Pedestrian Trajectory Prediction. In AAAI Conference on Artificial Intelligence. Liushuai Shi, Le Wang, Chengjiang Long, Sanping Zhou, Fang Zheng, Nanning Zheng, and Gang Hua. 2022. Social Interpretable Tree for Pedestrian Trajectory Prediction. In AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_2_2_43_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 8994--9003","author":"Shi Liushuai","year":"2021","unstructured":"Liushuai Shi , Le Wang , Chengjiang Long , Sanping Zhou , Mo Zhou , Zhenxing Niu , and Gang Hua . 2021 . SGCN: Sparse Graph Convolution for Pedestrian Trajectory Prediction . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 8994--9003 . Liushuai Shi, Le Wang, Chengjiang Long, Sanping Zhou, Mo Zhou, Zhenxing Niu, and Gang Hua. 2021. SGCN: Sparse Graph Convolution for Pedestrian Trajectory Prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 8994--9003."},{"key":"e_1_3_2_2_44_1","volume-title":"Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion. International journal of computer vision","author":"Sigal Leonid","year":"2010","unstructured":"Leonid Sigal , Alexandru O Balan , and Michael J Black . 2010 . Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion. International journal of computer vision , Vol. 87 , 1 (2010), 4--27. Leonid Sigal, Alexandru O Balan, and Michael J Black. 2010. Humaneva: Synchronized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion. International journal of computer vision, Vol. 87, 1 (2010), 4--27."},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v31i1.11212"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3474085.3475237"},{"key":"e_1_3_2_2_47_1","volume-title":"Intention-based Long-Term Human Motion Anticipation. In 2021 International Conference on 3D Vision (3DV). IEEE, 596--605","author":"Tanke Julian","year":"2021","unstructured":"Julian Tanke , Chintan Zaveri , and Juergen Gall . 2021 . Intention-based Long-Term Human Motion Anticipation. In 2021 International Conference on 3D Vision (3DV). IEEE, 596--605 . Julian Tanke, Chintan Zaveri, and Juergen Gall. 2021. Intention-based Long-Term Human Motion Anticipation. In 2021 International Conference on 3D Vision (3DV). IEEE, 596--605."},{"key":"e_1_3_2_2_48_1","volume-title":"Attention is all you need. Advances in neural information processing systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , \u0141ukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need. Advances in neural information processing systems , Vol. 30 ( 2017 ). Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, Vol. 30 (2017)."},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.361"},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01228-1_17"},{"key":"e_1_3_2_2_51_1","volume-title":"Diverse trajectory forecasting with determinantal point processes. arXiv preprint arXiv:1907.04967","author":"Yuan Ye","year":"2019","unstructured":"Ye Yuan and Kris Kitani . 2019. Diverse trajectory forecasting with determinantal point processes. arXiv preprint arXiv:1907.04967 ( 2019 ). Ye Yuan and Kris Kitani. 2019. Diverse trajectory forecasting with determinantal point processes. arXiv preprint arXiv:1907.04967 (2019)."},{"key":"e_1_3_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58545-7_20"}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3547956","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3547956","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:31Z","timestamp":1750186831000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3547956"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":52,"alternative-id":["10.1145\/3503161.3547956","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3547956","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}