{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,17]],"date-time":"2026-02-17T12:49:44Z","timestamp":1771332584274,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":46,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,14]]},"DOI":"10.1145\/3552437.3555693","type":"proceedings-article","created":{"date-parts":[[2022,9,30]],"date-time":"2022-09-30T22:08:38Z","timestamp":1664575718000},"page":"103-109","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["A Transformer-based System for Action Spotting in Soccer Videos"],"prefix":"10.1145","author":[{"given":"He","family":"Zhu","sequence":"first","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junwei","family":"Liang","sequence":"additional","affiliation":[{"name":"Tencent Youtu Lab, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengzhi","family":"Lin","sequence":"additional","affiliation":[{"name":"Sun Yat-sen University, Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Zhang","sequence":"additional","affiliation":[{"name":"Tencent Youtu Lab, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianming","family":"Hu","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58604-1_8"},{"key":"e_1_3_2_2_2_1","volume-title":"Procedings of the British Machine Vision Conference","author":"Buch Shyamal","year":"2019","unstructured":"Shyamal Buch , Victor Escorcia , Bernard Ghanem , Li Fei-Fei , and Juan Carlos Niebles . 2019 . End-to-end, single-stream temporal action detection in untrimmed videos . In Procedings of the British Machine Vision Conference 2017. British Machine Vision Association. Shyamal Buch, Victor Escorcia, Bernard Ghanem, Li Fei-Fei, and Juan Carlos Niebles. 2019. End-to-end, single-stream temporal action detection in untrimmed videos. In Procedings of the British Machine Vision Conference 2017. British Machine Vision Association."},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.675"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"crossref","unstructured":"Jo a o Carreira and Andrew Zisserman. 2017. Quo Vadis Action Recognition? A New Model and the Kinetics Dataset. In CVPR.  Jo a o Carreira and Andrew Zisserman. 2017. Quo Vadis Action Recognition? A New Model and the Kinetics Dataset. In CVPR.","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_2_5_1","volume-title":"Augmented transformer with adaptive graph for temporal action proposal generation. arXiv preprint arXiv:2103.16024","author":"Chang Shuning","year":"2021","unstructured":"Shuning Chang , Pichao Wang , Fan Wang , Hao Li , and Jiashi Feng . 2021. Augmented transformer with adaptive graph for temporal action proposal generation. arXiv preprint arXiv:2103.16024 ( 2021 ). Shuning Chang, Pichao Wang, Fan Wang, Hao Li, and Jiashi Feng. 2021. Augmented transformer with adaptive graph for temporal action proposal generation. arXiv preprint arXiv:2103.16024 (2021)."},{"key":"e_1_3_2_2_6_1","unstructured":"Xiaojun Chang Wenhe Liu Po-Yao Huang Changlin Li Fengda Zhu Mingfei Han Mingjie Li Mengyuan Ma Siyi Hu Guoliang Kang etal 2019. MMVG-INF-Etrol@ TRECVID 2019: Activities in Extended Video. (2019).  Xiaojun Chang Wenhe Liu Po-Yao Huang Changlin Li Fengda Zhu Mingfei Han Mingjie Li Mengyuan Ma Siyi Hu Guoliang Kang et al. 2019. MMVG-INF-Etrol@ TRECVID 2019: Activities in Extended Video. (2019)."},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01314"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00511"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00508"},{"key":"e_1_3_2_2_10_1","volume-title":"International Conference on Learning Representations.","author":"Dosovitskiy Alexey","year":"2020","unstructured":"Alexey Dosovitskiy , Lucas Beyer , Alexander Kolesnikov , Dirk Weissenborn , Xiaohua Zhai , Thomas Unterthiner , Mostafa Dehghani , Matthias Minderer , Georg Heigold , Sylvain Gelly , 2020 . An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale . In International Conference on Learning Representations. Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations."},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.812758"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00675"},{"key":"e_1_3_2_2_13_1","volume-title":"Storage and Retrieval Methods and Applications for Multimedia","volume":"5307","author":"Farin Dirk","year":"2003","unstructured":"Dirk Farin , Susanne Krabbe , Wolfgang Effelsberg , 2003 . Robust camera calibration for sport videos using court models . In Storage and Retrieval Methods and Applications for Multimedia 2004, Vol. 5307 . SPIE, 80--91. Dirk Farin, Susanne Krabbe, Wolfgang Effelsberg, et al. 2003. Robust camera calibration for sport videos using court models. In Storage and Retrieval Methods and Applications for Multimedia 2004, Vol. 5307. SPIE, 80--91."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"crossref","unstructured":"Christoph Feichtenhofer Haoqi Fan Jitendra Malik and Kaiming He. 2019. Slowfast networks for video recognition. In ICCV.  Christoph Feichtenhofer Haoqi Fan Jitendra Malik and Kaiming He. 2019. Slowfast networks for video recognition. In ICCV.","DOI":"10.1109\/ICCV.2019.00630"},{"key":"e_1_3_2_2_15_1","volume-title":"Wildes","author":"Feichtenhofer Christoph","year":"2016","unstructured":"Christoph Feichtenhofer , Axel Pinz , and Richard P . Wildes . 2016 . Spatiotemporal Residual Networks for Video Action Recognition. In NeurIPS, , Daniel D. Lee, Masashi Sugiyama, Ulrike von Luxburg, Isabelle Guyon, and Roman Garnett (Eds .). Christoph Feichtenhofer, Axel Pinz, and Richard P. Wildes. 2016. Spatiotemporal Residual Networks for Video Action Recognition. In NeurIPS, , Daniel D. Lee, Masashi Sugiyama, Ulrike von Luxburg, Isabelle Guyon, and Roman Garnett (Eds.)."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00506"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME46284.2020.9102850"},{"key":"e_1_3_2_2_18_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR.  Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.427"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2006.876289"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_3_2_2_22_1","unstructured":"Po-Yao Huang Junwei Liang Vaibhav Vaibhav Xiaojun Chang and Alexander Hauptmann. 2018. Informedia@ TRECVID 2018: Ad-hoc video search with discrete and continuous representations.  Po-Yao Huang Junwei Liang Vaibhav Vaibhav Xiaojun Chang and Alexander Hauptmann. 2018. Informedia@ TRECVID 2018: Ad-hoc video search with discrete and continuous representations."},{"key":"e_1_3_2_2_23_1","unstructured":"Will Kay Joao Carreira Karen Simonyan Brian Zhang Chloe Hillier Sudheendra Vijayanarasimhan Fabio Viola Tim Green Trevor Back Paul Natsev etal 2017. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950 (2017).  Will Kay Joao Carreira Karen Simonyan Brian Zhang Chloe Hillier Sudheendra Vijayanarasimhan Fabio Viola Tim Green Trevor Back Paul Natsev et al. 2017. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950 (2017)."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"crossref","unstructured":"Alexander Kl\"a ser Marcin Marszalek and Cordelia Schmid. 2008. A Spatio-Temporal Descriptor Based on 3D-Gradients. In BMVC Mark Everingham Chris J. Needham and Roberto Fraile (Eds.).  Alexander Kl\"a ser Marcin Marszalek and Cordelia Schmid. 2008. A Spatio-Temporal Descriptor Based on 3D-Gradients. In BMVC Mark Everingham Chris J. Needham and Roberto Fraile (Eds.).","DOI":"10.5244\/C.22.99"},{"key":"e_1_3_2_2_25_1","volume-title":"Improved multiscale vision transformers for classification and detection. arXiv preprint arXiv:2112.01526","author":"Li Yanghao","year":"2021","unstructured":"Yanghao Li , Chao-Yuan Wu , Haoqi Fan , Karttikeya Mangalam , Bo Xiong , Jitendra Malik , and Christoph Feichtenhofer . 2021. Improved multiscale vision transformers for classification and detection. arXiv preprint arXiv:2112.01526 ( 2021 ). Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. 2021. Improved multiscale vision transformers for classification and detection. arXiv preprint arXiv:2112.01526 (2021)."},{"key":"e_1_3_2_2_26_1","volume-title":"Spatial-Temporal Alignment Network for Action Recognition and Detection. arXiv preprint arXiv:2012.02426","author":"Liang Junwei","year":"2020","unstructured":"Junwei Liang , Liangliang Cao , Xuehan Xiong , Ting Yu , and Alexander Hauptmann . 2020. Spatial-Temporal Alignment Network for Action Recognition and Detection. arXiv preprint arXiv:2012.02426 ( 2020 ). Junwei Liang, Liangliang Cao, Xuehan Xiong, Ting Yu, and Alexander Hauptmann. 2020. Spatial-Temporal Alignment Network for Action Recognition and Detection. arXiv preprint arXiv:2012.02426 (2020)."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952426"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3078971.3079003"},{"key":"e_1_3_2_2_29_1","unstructured":"Junwei Liang Lu Jiang Deyu Meng and Alexander G Hauptmann. 2016. Learning to Detect Concepts from Webly-Labeled Video Data. In IJCAI.  Junwei Liang Lu Jiang Deyu Meng and Alexander G Hauptmann. 2016. Learning to Detect Concepts from Webly-Labeled Video Data. In IJCAI."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00356"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123343"},{"key":"e_1_3_2_2_32_1","volume-title":"Argus: Efficient activity detection system for extended video analysis. In WACVW.","author":"Liu Wenhe","year":"2020","unstructured":"Wenhe Liu , Guoliang Kang , Po-Yao Huang , Xiaojun Chang , Yijun Qian , Junwei Liang , Liangke Gui , Jing Wen , and Peng Chen . 2020 . Argus: Efficient activity detection system for extended video analysis. In WACVW. Wenhe Liu, Guoliang Kang, Po-Yao Huang, Xiaojun Chang, Yijun Qian, Junwei Liang, Liangke Gui, Jing Wen, and Peng Chen. 2020. Argus: Efficient activity detection system for extended video analysis. In WACVW."},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00043"},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"crossref","unstructured":"Thomas B Moeslund Graham Thomas Adrian Hilton etal 2014. Computer vision in sports. Vol. 1. Springer.  Thomas B Moeslund Graham Thomas Adrian Hilton et al. 2014. Computer vision in sports. Vol. 1. Springer.","DOI":"10.1007\/978-3-319-09396-3_1"},{"key":"e_1_3_2_2_35_1","volume-title":"Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfreund, Carl Vondrick, et al.","author":"Monfort Mathew","year":"2019","unstructured":"Mathew Monfort , Alex Andonian , Bolei Zhou , Kandan Ramakrishnan , Sarah Adel Bargal , Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfreund, Carl Vondrick, et al. 2019 . Moments in time dataset: one million videos for event understanding. IEEE transactions on pattern analysis and machine intelligence, Vol. 42 , 2 (2019), 502--508. Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfreund, Carl Vondrick, et al. 2019. Moments in time dataset: one million videos for event understanding. IEEE transactions on pattern analysis and machine intelligence, Vol. 42, 2 (2019), 502--508."},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2019.00307"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.599"},{"key":"e_1_3_2_2_38_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"crossref","unstructured":"Graham W. Taylor Rob Fergus Yann LeCun and Christoph Bregler. 2010. Convolutional Learning of Spatio-temporal Features. In ECCV Kostas Daniilidis Petros Maragos and Nikos Paragios (Eds.).  Graham W. Taylor Rob Fergus Yann LeCun and Christoph Bregler. 2010. Convolutional Learning of Spatio-temporal Features. In ECCV Kostas Daniilidis Petros Maragos and Nikos Paragios (Eds.).","DOI":"10.1007\/978-3-642-15567-3_11"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2018.00227"},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"crossref","unstructured":"Du Tran Lubomir D. Bourdev Rob Fergus Lorenzo Torresani and Manohar Paluri. 2015. Learning Spatiotemporal Features with 3D Convolutional Networks. In ICCV.  Du Tran Lubomir D. Bourdev Rob Fergus Lorenzo Torresani and Manohar Paluri. 2015. Learning Spatiotemporal Features with 3D Convolutional Networks. In ICCV.","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW50498.2020.00456"},{"key":"e_1_3_2_2_43_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In NIPS.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In NIPS."},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3016486"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00719"},{"key":"e_1_3_2_2_46_1","volume-title":"Feature Combination Meets Attention: Baidu Soccer Embeddings and Transformer based Temporal Detection. arXiv preprint arXiv:2106.14447","author":"Zhou Xin","year":"2021","unstructured":"Xin Zhou , Le Kang , Zhiyu Cheng , Bo He , and Jingyu Xin . 2021. Feature Combination Meets Attention: Baidu Soccer Embeddings and Transformer based Temporal Detection. arXiv preprint arXiv:2106.14447 ( 2021 ). io Xin Zhou, Le Kang, Zhiyu Cheng, Bo He, and Jingyu Xin. 2021. Feature Combination Meets Attention: Baidu Soccer Embeddings and Transformer based Temporal Detection. arXiv preprint arXiv:2106.14447 (2021). io"}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 5th International ACM Workshop on Multimedia Content Analysis in Sports"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3552437.3555693","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3552437.3555693","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:47:41Z","timestamp":1750178861000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3552437.3555693"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":46,"alternative-id":["10.1145\/3552437.3555693","10.1145\/3552437"],"URL":"https:\/\/doi.org\/10.1145\/3552437.3555693","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}