{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,8]],"date-time":"2025-09-08T05:31:02Z","timestamp":1757309462766,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":36,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T00:00:00Z","timestamp":1602460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,12]]},"DOI":"10.1145\/3394171.3416299","type":"proceedings-article","created":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T12:26:25Z","timestamp":1602505585000},"page":"4630-4634","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["Towards Accurate Human Pose Estimation in Videos of Crowded Scenes"],"prefix":"10.1145","author":[{"given":"Shuning","family":"Chang","sequence":"first","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Li","family":"Yuan","sequence":"additional","affiliation":[{"name":"National Univerisity of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuecheng","family":"Nie","sequence":"additional","affiliation":[{"name":"YiTu Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ziyuan","family":"Huang","sequence":"additional","affiliation":[{"name":"National Univerisity of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yichen","family":"Zhou","sequence":"additional","affiliation":[{"name":"YiTu Technology, Singapore, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yupeng","family":"Chen","sequence":"additional","affiliation":[{"name":"YiTu Technology, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiashi","family":"Feng","sequence":"additional","affiliation":[{"name":"National University of Singapore, Beijing, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuicheng","family":"Yan","sequence":"additional","affiliation":[{"name":"YiTu Technology, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,10,12]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00644"},{"key":"e_1_3_2_2_2_1","volume-title":"A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987","author":"Carreira Joao","year":"2019","unstructured":"Joao Carreira , Eric Noland , Chloe Hillier , and Andrew Zisserman . 2019. A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987 ( 2019 ). Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. 2019. A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987 (2019)."},{"key":"e_1_3_2_2_3_1","volume-title":"Multiple Predictions. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 12214--12223","author":"Chu Xuangeng","year":"2020","unstructured":"Xuangeng Chu , Anlin Zheng , Xiangyu Zhang , and Jian Sun . 2020 . Detection in Crowded Scenes: One Proposal , Multiple Predictions. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 12214--12223 . Xuangeng Chu, Anlin Zheng, Xiangyu Zhang, and Jian Sun. 2020. Detection in Crowded Scenes: One Proposal, Multiple Predictions. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 12214--12223."},{"key":"e_1_3_2_2_4_1","volume-title":"Pedestrian detection: An evaluation of the state of the art","author":"Dollar Piotr","year":"2011","unstructured":"Piotr Dollar , Christian Wojek , Bernt Schiele , and Pietro Perona . 2011. Pedestrian detection: An evaluation of the state of the art . IEEE transactions on pattern analysis and machine intelligence, Vol. 34 , 4 ( 2011 ), 743--761. Piotr Dollar, Christian Wojek, Bernt Schiele, and Pietro Perona. 2011. Pedestrian detection: An evaluation of the state of the art. IEEE transactions on pattern analysis and machine intelligence, Vol. 34, 4 (2011), 743--761."},{"key":"e_1_3_2_2_5_1","unstructured":"Christoph Feichtenhofer Haoqi Fan Jitendra Malik and Kaiming He. [n.d.]. SlowFast Networks for Video Recognition. ([n. d.]).  Christoph Feichtenhofer Haoqi Fan Jitendra Malik and Kaiming He. [n.d.]. SlowFast Networks for Video Recognition. ([n. d.])."},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00630"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.169"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.81"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00633"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.322"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00745"},{"key":"e_1_3_2_2_13_1","volume-title":"Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy . 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167 ( 2015 ). Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167 (2015)."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.106"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.106"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.324"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_2_18_1","volume-title":"2020 a. Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events. arXiv preprint arXiv:2005.04490","author":"Lin Weiyao","year":"2020","unstructured":"Weiyao Lin , Huabin Liu , Shizhan Liu , Yuxi Li , Guo-Jun Qi , Rui Qian , Tao Wang , Nicu Sebe , Ning Xu , Hongkai Xiong , 2020 a. Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events. arXiv preprint arXiv:2005.04490 ( 2020 ). Weiyao Lin, Huabin Liu, Shizhan Liu, Yuxi Li, Guo-Jun Qi, Rui Qian, Tao Wang, Nicu Sebe, Ning Xu, Hongkai Xiong, et al. 2020 a. Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events. arXiv preprint arXiv:2005.04490 (2020)."},{"key":"e_1_3_2_2_19_1","volume-title":"2020 b. Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events. arXiv preprint arXiv:2005.04490","author":"Lin Weiyao","year":"2020","unstructured":"Weiyao Lin , Huabin Liu , Shizhan Liu , Yuxi Li , Guo-Jun Qi , Rui Qian , Tao Wang , Nicu Sebe , Ning Xu , Hongkai Xiong , 2020 b. Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events. arXiv preprint arXiv:2005.04490 ( 2020 ). Weiyao Lin, Huabin Liu, Shizhan Liu, Yuxi Li, Guo-Jun Qi, Rui Qian, Tao Wang, Nicu Sebe, Ning Xu, Hongkai Xiong, et al. 2020 b. Human in Events: A Large-Scale Benchmark for Human-centric Video Analysis in Complex Events. arXiv preprint arXiv:2005.04490 (2020)."},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.91"},{"key":"e_1_3_2_2_22_1","volume-title":"Faster, Stronger. arXiv preprint arXiv:1612.08242","author":"Redmon Joseph","year":"2016","unstructured":"Joseph Redmon and Ali Farhadi . 2016. YOLO9000 : Better , Faster, Stronger. arXiv preprint arXiv:1612.08242 ( 2016 ). Joseph Redmon and Ali Farhadi. 2016. YOLO9000: Better, Faster, Stronger. arXiv preprint arXiv:1612.08242 (2016)."},{"key":"e_1_3_2_2_23_1","unstructured":"Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems. 91--99.  Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems. 91--99."},{"key":"e_1_3_2_2_24_1","volume-title":"Crowdhuman: A benchmark for detecting human in a crowd. arXiv preprint arXiv:1805.00123","author":"Shao Shuai","year":"2018","unstructured":"Shuai Shao , Zijian Zhao , Boxun Li , Tete Xiao , Gang Yu , Xiangyu Zhang , and Jian Sun . 2018 . Crowdhuman: A benchmark for detecting human in a crowd. arXiv preprint arXiv:1805.00123 (2018). Shuai Shao, Zijian Zhao, Boxun Li, Tete Xiao, Gang Yu, Xiangyu Zhang, and Jian Sun. 2018. Crowdhuman: A benchmark for detecting human in a crowd. arXiv preprint arXiv:1805.00123 (2018)."},{"key":"e_1_3_2_2_25_1","volume-title":"Asynchronous Interaction Aggregation for Action Detection. arXiv preprint arXiv:2004.07485","author":"Tang Jiajun","year":"2020","unstructured":"Jiajun Tang , Jin Xia , Xinzhi Mu , Bo Pang , and Cewu Lu. 2020. Asynchronous Interaction Aggregation for Action Detection. arXiv preprint arXiv:2004.07485 ( 2020 ). Jiajun Tang, Jin Xia, Xinzhi Mu, Bo Pang, and Cewu Lu. 2020. Asynchronous Interaction Aggregation for Action Detection. arXiv preprint arXiv:2004.07485 (2020)."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-013-0664-6"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00507"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00734"},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.634"},{"key":"e_1_3_2_2_30_1","volume-title":"Guilin Li, Tao Wang, and Jiashi Feng. 2019 a. Revisit knowledge distillation: a teacher-free framework. arXiv preprint arXiv:1909.11723","author":"Yuan Li","year":"2019","unstructured":"Li Yuan , Francis EH Tay , Guilin Li, Tao Wang, and Jiashi Feng. 2019 a. Revisit knowledge distillation: a teacher-free framework. arXiv preprint arXiv:1909.11723 ( 2019 ). Li Yuan, Francis EH Tay, Guilin Li, Tao Wang, and Jiashi Feng. 2019 a. Revisit knowledge distillation: a teacher-free framework. arXiv preprint arXiv:1909.11723 (2019)."},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33019143"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00315"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.474"},{"key":"e_1_3_2_2_34_1","volume-title":"Semantic understanding of scenes through the ade20k dataset. arXiv preprint arXiv:1608.05442","author":"Zhou Bolei","year":"2016","unstructured":"Bolei Zhou , Hang Zhao , Xavier Puig , Sanja Fidler , Adela Barriuso , and Antonio Torralba . 2016. Semantic understanding of scenes through the ade20k dataset. arXiv preprint arXiv:1608.05442 ( 2016 ). Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. 2016. Semantic understanding of scenes through the ade20k dataset. arXiv preprint arXiv:1608.05442 (2016)."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.544"},{"key":"e_1_3_2_2_36_1","volume-title":"Object Relation Detection Based on One-shot Learning. arXiv preprint arXiv:1807.05857","author":"Zhou Li","year":"2018","unstructured":"Li Zhou , Jian Zhao , Jianshu Li , Li Yuan , and Jiashi Feng . 2018. Object Relation Detection Based on One-shot Learning. arXiv preprint arXiv:1807.05857 ( 2018 ). Li Zhou, Jian Zhao, Jianshu Li, Li Yuan, and Jiashi Feng. 2018. Object Relation Detection Based on One-shot Learning. arXiv preprint arXiv:1807.05857 (2018)."}],"event":{"name":"MM '20: The 28th ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Seattle WA USA","acronym":"MM '20"},"container-title":["Proceedings of the 28th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3416299","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3394171.3416299","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:01:40Z","timestamp":1750197700000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3416299"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,12]]},"references-count":36,"alternative-id":["10.1145\/3394171.3416299","10.1145\/3394171"],"URL":"https:\/\/doi.org\/10.1145\/3394171.3416299","relation":{},"subject":[],"published":{"date-parts":[[2020,10,12]]},"assertion":[{"value":"2020-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}