{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T05:17:12Z","timestamp":1784179032580,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":70,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3552458.3556443","type":"proceedings-article","created":{"date-parts":[[2022,10,4]],"date-time":"2022-10-04T22:08:06Z","timestamp":1664921286000},"page":"41-50","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":26,"title":["Augmented Transformer with Adaptive Graph for Temporal Action Proposal Generation"],"prefix":"10.1145","author":[{"given":"Shuning","family":"Chang","sequence":"first","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Pichao","family":"Wang","sequence":"additional","affiliation":[{"name":"Alibaba Group, Seattle, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fan","family":"Wang","sequence":"additional","affiliation":[{"name":"Alibaba Group, Seattle, WA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hao","family":"Li","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zheng","family":"Shou","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Vivit: A video vision transformer. arXiv preprint arXiv:2103.15691","author":"Arnab Anurag","year":"2021","unstructured":"Anurag Arnab , Mostafa Dehghani , Georg Heigold , Chen Sun , Mario Lui , and Cordelia Schmid . 2021 . Vivit: A video vision transformer. arXiv preprint arXiv:2103.15691 (2021). Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lui, and Cordelia Schmid. 2021. Vivit: A video vision transformer. arXiv preprint arXiv:2103.15691 (2021)."},{"key":"e_1_3_2_2_2_1","volume-title":"Boundary Content Graph Neural Network for Temporal Action Proposal Generation. arXiv preprint arXiv:2008.01432","author":"Bai Yueran","year":"2020","unstructured":"Yueran Bai , Yingying Wang , Yunhai Tong , Yang Yang , Qiyue Liu , and Junhui Liu . 2020. Boundary Content Graph Neural Network for Temporal Action Proposal Generation. arXiv preprint arXiv:2008.01432 ( 2020 ). Yueran Bai, Yingying Wang, Yunhai Tong, Yang Yang, Qiyue Liu, and Junhui Liu. 2020. Boundary Content Graph Neural Network for Temporal Action Proposal Generation. arXiv preprint arXiv:2008.01432 (2020)."},{"key":"e_1_3_2_2_3_1","volume-title":"SST: Single-Stream Temporal Action Proposals. In CVPR.","author":"Buch Shyamal","year":"2017","unstructured":"Shyamal Buch , Victor Escorcia , Chuanqi Shen , Bernard Ghanem , and Juan Carlos Niebles . 2017 . SST: Single-Stream Temporal Action Proposals. In CVPR. Shyamal Buch, Victor Escorcia, Chuanqi Shen, Bernard Ghanem, and Juan Carlos Niebles. 2017. SST: Single-Stream Temporal Action Proposals. In CVPR."},{"key":"e_1_3_2_2_4_1","volume-title":"Activitynet: A large-scale video benchmark for human activity understanding. In CVPR.","author":"Heilbron Fabian Caba","year":"2015","unstructured":"Fabian Caba Heilbron , Victor Escorcia , Bernard Ghanem , and Juan Carlos Niebles . 2015 . Activitynet: A large-scale video benchmark for human activity understanding. In CVPR. Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015. Activitynet: A large-scale video benchmark for human activity understanding. In CVPR."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"crossref","unstructured":"Nicolas Carion Francisco Massa Gabriel Synnaeve Nicolas Usunier Alexander Kirillov and Sergey Zagoruyko. 2020. End-to-End Object Detection with Transformers. In ECCV.  Nicolas Carion Francisco Massa Gabriel Synnaeve Nicolas Usunier Alexander Kirillov and Sergey Zagoruyko. 2020. End-to-End Object Detection with Transformers. In ECCV.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00124"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2959977"},{"key":"e_1_3_2_2_9_1","volume-title":"International conference on machine learning. PMLR, 933--941","author":"Dauphin Yann N","year":"2017","unstructured":"Yann N Dauphin , Angela Fan , Michael Auli , and David Grangier . 2017 . Language modeling with gated convolutional networks . In International conference on machine learning. PMLR, 933--941 . Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017. Language modeling with gated convolutional networks. In International conference on machine learning. PMLR, 933--941."},{"key":"e_1_3_2_2_10_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_2_11_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly etal 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020).  Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00028"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00630"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"crossref","unstructured":"Christoph Feichtenhofer Axel Pinz and Andrew Zisserman. 2016. Convolutional two-stream network fusion for video action recognition. In CVPR.  Christoph Feichtenhofer Axel Pinz and Andrew Zisserman. 2016. Convolutional two-stream network fusion for video action recognition. In CVPR.","DOI":"10.1109\/CVPR.2016.213"},{"key":"e_1_3_2_2_15_1","volume-title":"CTAP: Complementary Temporal Action Proposal Generation. In ECCV.","author":"Gao Jiyang","year":"2018","unstructured":"Jiyang Gao , Kan Chen , and Ram Nevatia . 2018 . CTAP: Complementary Temporal Action Proposal Generation. In ECCV. Jiyang Gao, Kan Chen, and Ram Nevatia. 2018. CTAP: Complementary Temporal Action Proposal Generation. In ECCV."},{"key":"e_1_3_2_2_16_1","volume-title":"Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action Localization. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Gao Junyu","year":"2022","unstructured":"Junyu Gao , Mengyuan Chen , and Changsheng Xu . 2022 . Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action Localization. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Junyu Gao, Mengyuan Chen, and Changsheng Xu. 2022. Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action Localization. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_2_17_1","volume-title":"TURN TAP: Temporal Unit Regression Network for Temporal Action Proposals. In ICCV.","author":"Gao Jiyang","year":"2017","unstructured":"Jiyang Gao , Zhenheng Yang , Kan Chen , Chen Sun , and Ram Nevatia . 2017 . TURN TAP: Temporal Unit Regression Network for Temporal Action Proposals. In ICCV. Jiyang Gao, Zhenheng Yang, Kan Chen, Chen Sun, and Ram Nevatia. 2017. TURN TAP: Temporal Unit Regression Network for Temporal Action Proposals. In ICCV."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00033"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2839534"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"crossref","unstructured":"Liang Han PichaoWang Zhaozheng Yin FanWang and Hao Li. 2020. Exploiting Better Feature Aggregation for Video Object Detection. In ACM MM.  Liang Han PichaoWang Zhaozheng Yin FanWang and Hao Li. 2020. Exploiting Better Feature Aggregation for Video Object Detection. In ACM MM.","DOI":"10.1145\/3394171.3413927"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"crossref","unstructured":"Jingwei Ji Kaidi Cao and Juan Carlos Niebles. 2019. Learning temporal action proposals with fewer labels. In ICCV. 7073--7082.  Jingwei Ji Kaidi Cao and Juan Carlos Niebles. 2019. Learning temporal action proposals with fewer labels. In ICCV. 7073--7082.","DOI":"10.1109\/ICCV.2019.00717"},{"key":"e_1_3_2_2_22_1","unstructured":"Yu Gang Jiang Jingen Liu A Roshan Zamir George Toderici Ivan Laptev Mubarak Shah and Rahul Sukthankar. 2014. THUMOS challenge: Action recognition with a large number of classes.  Yu Gang Jiang Jingen Liu A Roshan Zamir George Toderici Ivan Laptev Mubarak Shah and Rahul Sukthankar. 2014. THUMOS challenge: Action recognition with a large number of classes."},{"key":"e_1_3_2_2_23_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Jimmy Ba . 2014 . Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014). Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"crossref","unstructured":"Alexander Kozlov Vadim Andronov and Yana Gritsenko. 2020. Lightweight network architecture for real-time action recognition. In SAC.  Alexander Kozlov Vadim Andronov and Yana Gritsenko. 2020. Lightweight network architecture for real-time action recognition. In SAC.","DOI":"10.1145\/3341105.3373906"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.2965434"},{"key":"e_1_3_2_2_26_1","unstructured":"Chuming Lin Jian Li Yabiao Wang Ying Tai Donghao Luo Zhipeng Cui Chengjie Wang Jilin Li Feiyue Huang and Rongrong Ji. 2020. Fast Learning of Temporal Action Proposal via Dense Boundary Generator.. In AAAI.  Chuming Lin Jian Li Yabiao Wang Ying Tai Donghao Luo Zhipeng Cui Chengjie Wang Jilin Li Feiyue Huang and Rongrong Ji. 2020. Fast Learning of Temporal Action Proposal via Dense Boundary Generator.. In AAAI."},{"key":"e_1_3_2_2_27_1","unstructured":"Chuming Lin Chengming Xu Donghao Luo Yabiao Wang Ying Tai Chengjie Wang Jilin Li Feiyue Huang and Yanwei Fu. 2021. Learning Salient Boundary Feature for Anchor-free Temporal Action Localization. In CVPR. 3320--3329.  Chuming Lin Chengming Xu Donghao Luo Yabiao Wang Ying Tai Chengjie Wang Jilin Li Feiyue Huang and Yanwei Fu. 2021. Learning Salient Boundary Feature for Anchor-free Temporal Action Localization. In CVPR. 3320--3329."},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00399"},{"key":"e_1_3_2_2_29_1","volume-title":"Bsn: Boundary sensitive network for temporal action proposal generation. In ECCV.","author":"Lin Tianwei","year":"2018","unstructured":"Tianwei Lin , Xu Zhao , Haisheng Su , ChongjingWang, and Ming Yang . 2018 . Bsn: Boundary sensitive network for temporal action proposal generation. In ECCV. Tianwei Lin, Xu Zhao, Haisheng Su, ChongjingWang, and Ming Yang. 2018. Bsn: Boundary sensitive network for temporal action proposal generation. In ECCV."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"crossref","unstructured":"Yuan Liu Lin Ma Yifeng Zhang Wei Liu and Shih-Fu Chang. 2019. Multi- Granularity Generator for Temporal Action Proposal. In CVPR.  Yuan Liu Lin Ma Yifeng Zhang Wei Liu and Shih-Fu Chang. 2019. Multi- Granularity Generator for Temporal Action Proposal. In CVPR.","DOI":"10.1109\/CVPR.2019.00372"},{"key":"e_1_3_2_2_31_1","volume-title":"Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030","author":"Liu Ze","year":"2021","unstructured":"Ze Liu , Yutong Lin , Yue Cao , Han Hu , Yixuan Wei , Zheng Zhang , Stephen Lin , and Baining Guo . 2021. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030 ( 2021 ). Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030 (2021)."},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2943204"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58568-6_17"},{"key":"e_1_3_2_2_34_1","volume-title":"The lear submission at thumos","author":"Oneata Dan","year":"2014","unstructured":"Dan Oneata , Jakob Verbeek , and Cordelia Schmid . 2014. The lear submission at thumos 2014 . (2014). Dan Oneata, Jakob Verbeek, and Cordelia Schmid. 2014. The lear submission at thumos 2014. (2014)."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"crossref","unstructured":"Zhiwu Qing Haisheng Su Weihao Gan Dongliang Wang Wei Wu Xiang Wang Yu Qiao Junjie Yan Changxin Gao and Nong Sang. 2021. Temporal Context Aggregation Network for Temporal Action Proposal Refinement. In CVPR. 485-- 494.  Zhiwu Qing Haisheng Su Weihao Gan Dongliang Wang Wei Wu Xiang Wang Yu Qiao Junjie Yan Changxin Gao and Nong Sang. 2021. Temporal Context Aggregation Network for Temporal Action Proposal Refinement. In CVPR. 485-- 494.","DOI":"10.1109\/CVPR46437.2021.00055"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.590"},{"key":"e_1_3_2_2_37_1","unstructured":"Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In NIPS.  Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In NIPS."},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00109"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.119"},{"key":"e_1_3_2_2_40_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In NIPS.  Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In NIPS."},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"crossref","unstructured":"Waqas Sultani Chen Chen and Mubarak Shah. 2018. Real-world anomaly detection in surveillance videos. In CVPR.  Waqas Sultani Chen Chen and Mubarak Shah. 2018. Real-world anomaly detection in surveillance videos. In CVPR.","DOI":"10.1109\/CVPR.2018.00678"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"crossref","unstructured":"Chen Sun Sanketh Shetty Rahul Sukthankar and Ram Nevatia. 2015. Temporal localization of fine-grained actions in videos by domain transfer from web images. In ACM MM.  Chen Sun Sanketh Shetty Rahul Sukthankar and Ram Nevatia. 2015. Temporal localization of fine-grained actions in videos by domain transfer from web images. In ACM MM.","DOI":"10.1145\/2733373.2806226"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"crossref","unstructured":"Lin Sun Kui Jia Dit-Yan Yeung and Bertram E Shi. 2015. Human action recognition using factorized spatio-temporal convolutional networks. In ICCV. 4597-- 4605.  Lin Sun Kui Jia Dit-Yan Yeung and Bertram E Shi. 2015. Human action recognition using factorized spatio-temporal convolutional networks. In ICCV. 4597-- 4605.","DOI":"10.1109\/ICCV.2015.522"},{"key":"e_1_3_2_2_44_1","volume-title":"Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877","author":"Touvron Hugo","year":"2020","unstructured":"Hugo Touvron , Matthieu Cord , Matthijs Douze , Francisco Massa , Alexandre Sablayrolles , and Herv\u00e9 J\u00e9gou . 2020. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877 ( 2020 ). Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv\u00e9 J\u00e9gou. 2020. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877 (2020)."},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"crossref","unstructured":"Du Tran Lubomir Bourdev Rob Fergus Lorenzo Torresani and Manohar Paluri. 2015. Learning spatiotemporal features with 3d convolutional networks. In ICCV.  Du Tran Lubomir Bourdev Rob Fergus Lorenzo Torresani and Manohar Paluri. 2015. Learning spatiotemporal features with 3d convolutional networks. In ICCV.","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"crossref","unstructured":"Du Tran Heng Wang Lorenzo Torresani and Matt Feiszli. 2019. Video classification with channel-separated convolutional networks. In ICCV. 5552--5561.  Du Tran Heng Wang Lorenzo Torresani and Matt Feiszli. 2019. Video classification with channel-separated convolutional networks. In ICCV. 5552--5561.","DOI":"10.1109\/ICCV.2019.00565"},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"crossref","unstructured":"Du Tran Heng Wang Lorenzo Torresani Jamie Ray Yann LeCun and Manohar Paluri. 2018. A closer look at spatiotemporal convolutions for action recognition. In CVPR. 6450--6459.  Du Tran Heng Wang Lorenzo Torresani Jamie Ray Yann LeCun and Manohar Paluri. 2018. A closer look at spatiotemporal convolutions for action recognition. In CVPR. 6450--6459.","DOI":"10.1109\/CVPR.2018.00675"},{"key":"e_1_3_2_2_48_1","volume-title":"ukasz Kaiser, and Illia Polosukhin","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , ukasz Kaiser, and Illia Polosukhin . 2017 . Attention is all you need. In NIPS. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NIPS."},{"key":"e_1_3_2_2_49_1","volume-title":"Max-deeplab: End-to-end panoptic segmentation with mask transformers. In CVPR. 5463--5474.","author":"Wang Huiyu","year":"2021","unstructured":"Huiyu Wang , Yukun Zhu , Hartwig Adam , Alan Yuille , and Liang-Chieh Chen . 2021 . Max-deeplab: End-to-end panoptic segmentation with mask transformers. In CVPR. 5463--5474. Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. 2021. Max-deeplab: End-to-end panoptic segmentation with mask transformers. In CVPR. 5463--5474."},{"key":"e_1_3_2_2_50_1","volume-title":"Action recognition and detection by combining motion and appearance features. THUMOS14 Action Recognition Challenge","author":"Wang Limin","year":"2014","unstructured":"Limin Wang , Yu Qiao , and Xiaoou Tang . 2014. Action recognition and detection by combining motion and appearance features. THUMOS14 Action Recognition Challenge ( 2014 ). Limin Wang, Yu Qiao, and Xiaoou Tang. 2014. Action recognition and detection by combining motion and appearance features. THUMOS14 Action Recognition Challenge (2014)."},{"key":"e_1_3_2_2_51_1","doi-asserted-by":"crossref","unstructured":"LiminWang Yuanjun Xiong Dahua Lin and Luc Van Gool. 2017. UntrimmedNets for Weakly Supervised Action Recognition and Detection. In CVPR.  LiminWang Yuanjun Xiong Dahua Lin and Luc Van Gool. 2017. UntrimmedNets for Weakly Supervised Action Recognition and Detection. In CVPR.","DOI":"10.1109\/CVPR.2017.678"},{"key":"e_1_3_2_2_52_1","doi-asserted-by":"crossref","unstructured":"Limin Wang Yuanjun Xiong Zhe Wang Yu Qiao Dahua Lin Xiaoou Tang and Luc Van Gool. 2016. Temporal segment networks: Towards good practices for deep action recognition. In ECCV.  Limin Wang Yuanjun Xiong Zhe Wang Yu Qiao Dahua Lin Xiaoou Tang and Luc Van Gool. 2016. Temporal segment networks: Towards good practices for deep action recognition. In ECCV.","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"e_1_3_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2818329"},{"key":"e_1_3_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2018.04.007"},{"key":"e_1_3_2_2_55_1","volume-title":"KVT: k-NN Attention for Boosting Vision Transformers. arXiv preprint arXiv:2106.00515","author":"Wang Pichao","year":"2021","unstructured":"Pichao Wang , Xue Wang , Fan Wang , Ming Lin , Shuning Chang , Wen Xie , Hao Li , and Rong Jin . 2021. KVT: k-NN Attention for Boosting Vision Transformers. arXiv preprint arXiv:2106.00515 ( 2021 ). Pichao Wang, Xue Wang, Fan Wang, Ming Lin, Shuning Chang, Wen Xie, Hao Li, and Rong Jin. 2021. KVT: k-NN Attention for Boosting Vision Transformers. arXiv preprint arXiv:2106.00515 (2021)."},{"key":"e_1_3_2_2_56_1","volume-title":"Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv preprint arXiv:2102.12122","author":"Xie Enze","year":"2021","unstructured":"WenhaiWang, Enze Xie , Xiang Li , Deng-Ping Fan , Kaitao Song , Ding Liang , Tong Lu , Ping Luo , and Ling Shao . 2021. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv preprint arXiv:2102.12122 ( 2021 ). WenhaiWang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2021. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv preprint arXiv:2102.12122 (2021)."},{"key":"e_1_3_2_2_57_1","doi-asserted-by":"crossref","unstructured":"XiaolongWang Ross Girshick Abhinav Gupta and Kaiming He. 2018. Non-Local Neural Networks. In CVPR.  XiaolongWang Ross Girshick Abhinav Gupta and Kaiming He. 2018. Non-Local Neural Networks. In CVPR.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"e_1_3_2_2_58_1","doi-asserted-by":"crossref","unstructured":"Xiang Wang Shiwei Zhang Zhiwu Qing Yuanjie Shao Changxin Gao and Nong Sang. 2021. Self-supervised learning for semi-supervised temporal action proposal. In CVPR. 1905--1914.  Xiang Wang Shiwei Zhang Zhiwu Qing Yuanjie Shao Changxin Gao and Nong Sang. 2021. Self-supervised learning for semi-supervised temporal action proposal. In CVPR. 1905--1914.","DOI":"10.1109\/CVPR46437.2021.00194"},{"key":"e_1_3_2_2_59_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).","author":"KongWentao Bao Yu","year":"2022","unstructured":"Yu KongWentao Bao , Qi Yu . 2022 . OpenTAL: Towards Open Set Temporal Action Localization . In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Yu KongWentao Bao, Qi Yu. 2022. OpenTAL: Towards Open Set Temporal Action Localization. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_2_60_1","unstructured":"Chao-Yuan Wu Christoph Feichtenhofer Haoqi Fan Kaiming He Philipp Krahenbuhl and Ross Girshick. 2019. Long-Term Feature Banks for Detailed Video Understanding. In CVPR.  Chao-Yuan Wu Christoph Feichtenhofer Haoqi Fan Kaiming He Philipp Krahenbuhl and Ross Girshick. 2019. Long-Term Feature Banks for Detailed Video Understanding. In CVPR."},{"key":"e_1_3_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2953814"},{"key":"e_1_3_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"e_1_3_2_2_63_1","unstructured":"Mengmeng Xu Chen Zhao David S Rojas Ali Thabet and Bernard Ghanem. 2020. G-TAD: Sub-Graph Localization for Temporal Action Detection. In CVPR.  Mengmeng Xu Chen Zhao David S Rojas Ali Thabet and Bernard Ghanem. 2020. G-TAD: Sub-Graph Localization for Temporal Action Detection. In CVPR."},{"key":"e_1_3_2_2_64_1","volume-title":"Jiashi Feng, and Shuicheng Yan.","author":"Yuan Li","year":"2021","unstructured":"Li Yuan , Yunpeng Chen , Tao Wang , Weihao Yu , Yujun Shi , Zihang Jiang , Francis EH Tay , Jiashi Feng, and Shuicheng Yan. 2021 . Tokens-to-token vit: Training vision transformers from scratch on imagenet. arXiv preprint arXiv:2101.11986 (2021). Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. 2021. Tokens-to-token vit: Training vision transformers from scratch on imagenet. arXiv preprint arXiv:2101.11986 (2021)."},{"key":"e_1_3_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58539-6_3"},{"key":"e_1_3_2_2_66_1","volume-title":"Point transformer. arXiv preprint arXiv:2012.09164","author":"Zhao Hengshuang","year":"2020","unstructured":"Hengshuang Zhao , Li Jiang , Jiaya Jia , Philip Torr , and Vladlen Koltun . 2020. Point transformer. arXiv preprint arXiv:2012.09164 ( 2020 ). Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip Torr, and Vladlen Koltun. 2020. Point transformer. arXiv preprint arXiv:2012.09164 (2020)."},{"key":"e_1_3_2_2_67_1","doi-asserted-by":"crossref","unstructured":"Peisen Zhao Lingxi Xie Chen Ju Ya Zhang Yanfeng Wang and Qi Tian. 2020. Bottom-Up Temporal Action Localization with Mutual Regularization. In ECCV.  Peisen Zhao Lingxi Xie Chen Ju Ya Zhang Yanfeng Wang and Qi Tian. 2020. Bottom-Up Temporal Action Localization with Mutual Regularization. In ECCV.","DOI":"10.1007\/978-3-030-58598-3_32"},{"key":"e_1_3_2_2_68_1","doi-asserted-by":"crossref","unstructured":"Yue Zhao Yuanjun Xiong Limin Wang Zhirong Wu Xiaoou Tang and Dahua Lin. 2017. Temporal action detection with structured segment networks. In ICCV.  Yue Zhao Yuanjun Xiong Limin Wang Zhirong Wu Xiaoou Tang and Dahua Lin. 2017. Temporal action detection with structured segment networks. In ICCV.","DOI":"10.1109\/ICCV.2017.317"},{"key":"e_1_3_2_2_69_1","unstructured":"Y Zhao B Zhang Z Wu S Yang L Zhou S Yan L Wang Y Xiong D Lin Y Qiao etal 2017. Cuhk & ethz & siat submission to activitynet challenge 2017. arXiv preprint arXiv:1710.08011 8 (2017).  Y Zhao B Zhang Z Wu S Yang L Zhou S Yan L Wang Y Xiong D Lin Y Qiao et al. 2017. Cuhk & ethz & siat submission to activitynet challenge 2017. arXiv preprint arXiv:1710.08011 8 (2017)."},{"key":"e_1_3_2_2_70_1","volume-title":"Philip HS Torr, et al","author":"Zheng Sixiao","year":"2021","unstructured":"Sixiao Zheng , Jiachen Lu , Hengshuang Zhao , Xiatian Zhu , Zekun Luo , Yabiao Wang , Yanwei Fu , Jianfeng Feng , Tao Xiang , Philip HS Torr, et al . 2021 . Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In CVPR. 6881--6890. Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. 2021. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In CVPR. 6881--6890."}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 3rd International Workshop on Human-Centric Multimedia Analysis"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3552458.3556443","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3552458.3556443","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:47:42Z","timestamp":1750178862000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3552458.3556443"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":70,"alternative-id":["10.1145\/3552458.3556443","10.1145\/3552458"],"URL":"https:\/\/doi.org\/10.1145\/3552458.3556443","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}