{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T19:02:16Z","timestamp":1784314936711,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":46,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T00:00:00Z","timestamp":1634428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100003399","name":"Science and Technology Commission of Shanghai Municipality","doi-asserted-by":"publisher","award":["2021SHZDZX0103"],"award-info":[{"award-number":["2021SHZDZX0103"]}],"id":[{"id":"10.13039\/501100003399","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["82090052"],"award-info":[{"award-number":["82090052"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,17]]},"DOI":"10.1145\/3474085.3475438","type":"proceedings-article","created":{"date-parts":[[2021,10,18]],"date-time":"2021-10-18T04:59:18Z","timestamp":1634533158000},"page":"4902-4910","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":87,"title":["TSA-Net: Tube Self-Attention Network for Action Quality Assessment"],"prefix":"10.1145","author":[{"given":"Shunli","family":"Wang","sequence":"first","affiliation":[{"name":"Fudan University &amp; Engineering Research Center of AI and Robotics, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dingkang","family":"Yang","sequence":"additional","affiliation":[{"name":"Fudan University &amp; Ji Hua Laboratory, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peng","family":"Zhai","sequence":"additional","affiliation":[{"name":"Fudan University &amp; Jilin Provincial Key Laboratory of Intelligence Science and Engineering, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chixiao","family":"Chen","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lihua","family":"Zhang","sequence":"additional","affiliation":[{"name":"Ji Hua Laboratory &amp; Fudan University, Foshan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,10,17]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/3326943.3326976"},{"key":"e_1_3_2_2_3_1","volume-title":"Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171--4186","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171--4186 . Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 4171--4186."},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00634"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00805"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00407"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.256"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00028"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.213"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240550"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00378"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2913372"},{"key":"e_1_3_2_2_13_1","volume-title":"CCNet: Criss-Cross Attention for Semantic Segmentation","author":"Huang Zilong","year":"2020","unstructured":"Zilong Huang , Xinggang Wang , Chang Huang , Yunchao Wei , Lichao Huang , and Wenyu Liu . 2020. CCNet: Criss-Cross Attention for Semantic Segmentation . IEEE Transactions on Pattern Analysis and Machine Intelligence ( 2020 ), 1--1. Zilong Huang, Xinggang Wang, Chang Huang, Yunchao Wei, Lichao Huang, and Wenyu Liu. 2020. CCNet: Criss-Cross Attention for Semantic Segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence (2020), 1--1."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/3104322.3104386"},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.223"},{"key":"e_1_3_2_2_16_1","volume-title":"Article arXiv:1705.06950","author":"Kay Will","year":"2017","unstructured":"Will Kay , Joao Carreira , Karen Simonyan , Brian Zhang , Chloe Hillier , Sudheendra Vijayanarasimhan , Fabio Viola , Tim Green , Trevor Back , Paul Natsev , Mustafa Suleyman , and Andrew Zisserman . 2017. The Kinetics Human Action Video Dataset ., Article arXiv:1705.06950 ( 2017 ). arXiv:1705.06950 Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman. 2017. The Kinetics Human Action Video Dataset., Article arXiv:1705.06950 (2017). arXiv:1705.06950"},{"key":"e_1_3_2_2_17_1","volume-title":"Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations.","author":"Kingma Diederik","year":"2014","unstructured":"Diederik Kingma and Jimmy Ba . 2014 . Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations. Diederik Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00718"},{"key":"e_1_3_2_2_19_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ECCV).","author":"Lin Tsung-Yi","unstructured":"Tsung-Yi Lin , Michael Maire , Serge J. Belongie , Lubomir D. Bourdev , Ross B. Girshick , James Hays , Pietro Perona , Deva Ramanan , Piotr Doll\u00e1r , and C. Lawrence Zitnick . 2014. Microsoft COCO: Common Objects in Context . In Proceedings of the IEEE International Conference on Computer Vision (ECCV). Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. In Proceedings of the IEEE International Conference on Computer Vision (ECCV)."},{"key":"e_1_3_2_2_20_1","volume-title":"Chi Chiung Grace Chen, and Gregory D. Hager","author":"Malpani Anand","year":"2014","unstructured":"Anand Malpani , S. Swaroop Vedula , Chi Chiung Grace Chen, and Gregory D. Hager . 2014 . Pairwise Comparison-Based Objective Score for Automated Skill Assessment of Segments in a Surgical Task. In Information Processing in Computer- Assisted Interventions . 138--147. Anand Malpani, S. Swaroop Vedula, Chi Chiung Grace Chen, and Gregory D. Hager. 2014. Pairwise Comparison-Based Objective Score for Automated Skill Assessment of Segments in a Surgical Task. In Information Processing in Computer- Assisted Interventions. 138--147."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969033.2969073"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00643"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2017.16"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2019.00161"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00039"},{"key":"e_1_3_2_2_26_1","volume-title":"Advances in Neural Information Processing Systems Workshops (NIPSW).","author":"Paszke Adam","unstructured":"Adam Paszke , S. Gross , Soumith Chintala , Gregory Chanan , Edward Yang , Zachary Devito , Zeming Lin , Alban Desmaison , L. Antiga , and A. Lerer . 2017. Automatic differentiation in PyTorch . In Advances in Neural Information Processing Systems Workshops (NIPSW). Adam Paszke, S. Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary Devito, Zeming Lin, Alban Desmaison, L. Antiga, and A. Lerer. 2017. Automatic differentiation in PyTorch. In Advances in Neural Information Processing Systems Workshops (NIPSW)."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10599-4_36"},{"key":"e_1_3_2_2_28_1","unstructured":"A. Radford and Karthik Narasimhan. 2018. Improving Language Understanding by Generative Pre-Training.  A. Radford and Karthik Narasimhan. 2018. Improving Language Understanding by Generative Pre-Training."},{"key":"e_1_3_2_2_29_1","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2018. Language Models are Unsupervised Multitask Learners. (2018).  Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2018. Language Models are Unsupervised Multitask Learners. (2018)."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2577031"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00269"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.5555\/2968826.2968890"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00986"},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00675"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_2_2_37_1","volume-title":"Dynamical Regularity for Action Analysis. In British Machine Vision Virtual Conference (BMVC). 67","author":"Venkataraman Vinay","year":"2015","unstructured":"Vinay Venkataraman , Ioannis Vlachos , and Pavan Turaga . 2015 . Dynamical Regularity for Action Analysis. In British Machine Vision Virtual Conference (BMVC). 67 .1--67.12. Vinay Venkataraman, Ioannis Vlachos, and Pavan Turaga. 2015. Dynamical Regularity for Action Analysis. In British Machine Vision Virtual Conference (BMVC). 67.1--67.12."},{"key":"e_1_3_2_2_38_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR). 1328--1338","author":"Wang Qiang","unstructured":"Qiang Wang , Li Zhang , Luca Bertinetto , Weiming Hu , and Philip H.S. Torr . 2019. Fast Online Object Tracking and Segmentation: A Unifying Approach . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR). 1328--1338 . Qiang Wang, Li Zhang, Luca Bertinetto, Weiming Hu, and Philip H.S. Torr. 2019. Fast Online Object Tracking and Segmentation: A Unifying Approach. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR). 1328--1338."},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00813"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2927118"},{"key":"e_1_3_2_2_41_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 7444--7452","author":"Yan Sijie","year":"2018","unstructured":"Sijie Yan , Yuanjun Xiong , and Dahua Lin . 2018 . Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition . In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 7444--7452 . Sijie Yan, Yuanjun Xiong, and Dahua Lin. 2018. Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 7444--7452."},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.5555\/3327757.3327758"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58604-1_20"},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01240-3_17"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11548-016-1468-2"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11548-017-1600-y"}],"event":{"name":"MM '21: ACM Multimedia Conference","location":"Virtual Event China","acronym":"MM '21","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 29th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475438","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3474085.3475438","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:48:33Z","timestamp":1750193313000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475438"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,17]]},"references-count":46,"alternative-id":["10.1145\/3474085.3475438","10.1145\/3474085"],"URL":"https:\/\/doi.org\/10.1145\/3474085.3475438","relation":{},"subject":[],"published":{"date-parts":[[2021,10,17]]},"assertion":[{"value":"2021-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}