{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T05:04:55Z","timestamp":1750309495678,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":50,"publisher":"ACM","license":[{"start":{"date-parts":[[2024,10,28]],"date-time":"2024-10-28T00:00:00Z","timestamp":1730073600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,10,28]]},"DOI":"10.1145\/3664647.3681329","type":"proceedings-article","created":{"date-parts":[[2024,10,26]],"date-time":"2024-10-26T06:59:27Z","timestamp":1729925967000},"page":"8139-8148","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Label Text-aided Hierarchical Semantics Mining for Panoramic Activity Recognition"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3831-8893","authenticated-orcid":false,"given":"Tianshan","family":"Liu","sequence":"first","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0422-8454","authenticated-orcid":false,"given":"Kin-Man","family":"Lam","sequence":"additional","affiliation":[{"name":"The Hong Kong Polytechnic University, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5956-831X","authenticated-orcid":false,"given":"Bing-Kun","family":"Bao","sequence":"additional","affiliation":[{"name":"Nanjing University of Posts and Telecommunications &amp; Peng Cheng Laboratory, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,10,28]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"International Conference on Machine Learning (ICML)","volume":"2","author":"Bertasius Gedas","year":"2021","unstructured":"Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding?. In International Conference on Machine Learning (ICML), Vol. 2. 4."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_2_1","DOI":"10.1145\/3581783.3612435"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_3_1","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_1_4_1","volume-title":"International Conference on Learning Representations (ICLR).","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR)."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_5_1","DOI":"10.1007\/978-3-030-58545-7_11"},{"key":"e_1_3_2_1_6_1","volume-title":"Social Group and Activity Detection. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 20951--20960","author":"Ehsanpour Mahsa","year":"2022","unstructured":"Mahsa Ehsanpour, Fatemeh Saleh, Silvio Savarese, Ian Reid, and Hamid Rezatofighi. 2022. JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 20951--20960."},{"key":"e_1_3_2_1_7_1","volume-title":"Multiscale Vision Transformers. In 2021 IEEE\/CVF International Conference on Computer Vision (ICCV). 6804--6815","author":"Fan Haoqi","year":"2021","unstructured":"Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. 2021. Multiscale Vision Transformers. In 2021 IEEE\/CVF International Conference on Computer Vision (ICCV). 6804--6815."},{"key":"e_1_3_2_1_8_1","volume-title":"SlowFast Networks for Video Recognition. In 2019 IEEE\/CVF International Conference on Computer Vision (ICCV). 6201--6210","author":"Feichtenhofer Christoph","year":"2019","unstructured":"Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019. SlowFast Networks for Video Recognition. In 2019 IEEE\/CVF International Conference on Computer Vision (ICCV). 6201--6210."},{"key":"e_1_3_2_1_9_1","volume-title":"Pyramidclip: Hierarchical feature alignment for vision-language model pretraining. Advances in neural information processing systems","author":"Gao Yuting","year":"2022","unstructured":"Yuting Gao, Jinfeng Liu, Zihan Xu, Jun Zhang, Ke Li, Rongrong Ji, and Chunhua Shen. 2022. Pyramidclip: Hierarchical feature alignment for vision-language model pretraining. Advances in neural information processing systems, Vol. 35 (2022), 35959--35970."},{"volume-title":"Actor-Transformers for Group Activity Recognition. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 836--845","author":"Gavrilyuk Kirill","unstructured":"Kirill Gavrilyuk, Ryan Sanford, Mehrsan Javan, and Cees G. M. Snoek. 2020. Actor-Transformers for Group Activity Recognition. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 836--845.","key":"e_1_3_2_1_10_1"},{"key":"e_1_3_2_1_11_1","volume-title":"Dual-AI: Dual-path Actor Interaction Learning for Group Activity Recognition. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2980--2989","author":"Han Mingfei","year":"2022","unstructured":"Mingfei Han, David Junhao Zhang, Yali Wang, Rui Yan, Lina Yao, Xiaojun Chang, and Yu Qiao. 2022. Dual-AI: Dual-path Actor Interaction Learning for Group Activity Recognition. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2980--2989."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_12_1","DOI":"10.1007\/978-3-031-19772-7_15"},{"key":"e_1_3_2_1_13_1","volume-title":"Mask R-CNN. In 2017 IEEE International Conference on Computer Vision (ICCV). 2980--2988","author":"He Kaiming","year":"2017","unstructured":"Kaiming He, Georgia Gkioxari, Piotr Doll\u00e1r, and Ross Girshick. 2017. Mask R-CNN. In 2017 IEEE International Conference on Computer Vision (ICCV). 2980--2988."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_14_1","DOI":"10.1109\/CVPR.2016.217"},{"key":"e_1_3_2_1_15_1","volume-title":"STM: SpatioTemporal and Motion Encoding for Action Recognition. In 2019 IEEE\/CVF International Conference on Computer Vision (ICCV). 2000--2009","author":"Jiang Boyuan","year":"2019","unstructured":"Boyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu, and Junjie Yan. 2019. STM: SpatioTemporal and Motion Encoding for Action Recognition. In 2019 IEEE\/CVF International Conference on Computer Vision (ICCV). 2000--2009."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_16_1","DOI":"10.1007\/978-3-031-19833-5_7"},{"key":"e_1_3_2_1_17_1","volume-title":"International Conference on Learning Representations (ICLR). 1--11","author":"Kingma Diederik P","year":"2015","unstructured":"Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR). 1--11."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_18_1","DOI":"10.1007\/s11263-022-01594-9"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_19_1","DOI":"10.18653\/v1\/2020.emnlp-main.161"},{"key":"e_1_3_2_1_20_1","volume-title":"GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal Transformer. In 2021 IEEE\/CVF International Conference on Computer Vision (ICCV). 13648--13657","author":"Li Shuaicheng","year":"2021","unstructured":"Shuaicheng Li, Qianggang Cao, Lingbo Liu, Kunlin Yang, Shinan Liu, Jun Hou, and Shuai Yi. 2021. GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal Transformer. In 2021 IEEE\/CVF International Conference on Computer Vision (ICCV). 13648--13657."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_21_1","DOI":"10.1145\/3503161.3548341"},{"key":"e_1_3_2_1_22_1","volume-title":"Video Swin Transformer. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 3192--3201","author":"Liu Ze","year":"2022","unstructured":"Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. 2022. Video Swin Transformer. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 3192--3201."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_23_1","DOI":"10.1109\/TPAMI.2021.3070543"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_24_1","DOI":"10.3115\/v1\/D14-1162"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_25_1","DOI":"10.1007\/978-3-030-58452-8_5"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_26_1","DOI":"10.1145\/3394171.3413954"},{"key":"e_1_3_2_1_27_1","volume-title":"DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 18061--18070","author":"Rao Yongming","year":"2022","unstructured":"Yongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang, Zheng Zhu, Guan Huang, Jie Zhou, and Jiwen Lu. 2022. DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 18061--18070."},{"volume-title":"2013 IEEE Conference on Computer Vision and Pattern Recognition. 2730--2737","author":"Michael","unstructured":"Michael S. Ryoo and Larry Matthies. 2013. First-Person Activity Recognition: What Are They Doing to Me?. In 2013 IEEE Conference on Computer Vision and Pattern Recognition. 2730--2737.","key":"e_1_3_2_1_28_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_29_1","DOI":"10.1109\/TPAMI.2019.2942030"},{"key":"e_1_3_2_1_30_1","volume-title":"Two-stream convolutional networks for action recognition in videos. Advances in neural information processing systems","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. Advances in neural information processing systems, Vol. 27 (2014)."},{"key":"e_1_3_2_1_31_1","first-page":"3200","article-title":"Human Action Recognition From Various Data Modalities","volume":"45","author":"Sun Zehua","year":"2023","unstructured":"Zehua Sun, Qiuhong Ke, Hossein Rahmani, Mohammed Bennamoun, Gang Wang, and Jun Liu. 2023. Human Action Recognition From Various Data Modalities: A Review. IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 45, 3 (2023), 3200--3225.","journal-title":"A Review. IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_1_32_1","volume-title":"Rethinking the Inception Architecture for Computer Vision. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2818--2826","author":"Szegedy Christian","year":"2016","unstructured":"Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the Inception Architecture for Computer Vision. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2818--2826."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_33_1","DOI":"10.1007\/978-3-031-19772-7_2"},{"key":"e_1_3_2_1_34_1","volume-title":"Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding. In 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 18973--18982","author":"Tan Chaolei","year":"2023","unstructured":"Chaolei Tan, Zihang Lin, Jian-Fang Hu, Wei-Shi Zheng, and Jianhuang Lai. 2023. Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding. In 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 18973--18982."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_35_1","DOI":"10.1109\/ICCV.2015.510"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_36_1","DOI":"10.1109\/CVPR.2018.00675"},{"key":"e_1_3_2_1_37_1","volume-title":"Attention is all you need. Advances in neural information processing systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, Vol. 30 (2017)."},{"key":"e_1_3_2_1_38_1","volume-title":"Deformable Video Transformer. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 14033--14042","author":"Wang Jue","year":"2022","unstructured":"Jue Wang and Lorenzo Torresani. 2022. Deformable Video Transformer. In 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 14033--14042."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_39_1","DOI":"10.1109\/TPAMI.2018.2868668"},{"key":"e_1_3_2_1_40_1","volume-title":"PANDA: A Gigapixel-Level Human-Centric Video Dataset. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 3265--3275","author":"Wang Xueyang","year":"2020","unstructured":"Xueyang Wang, Xiya Zhang, Yinheng Zhu, Yuchen Guo, Xiaoyun Yuan, Liuyu Xiang, Zerun Wang, Guiguang Ding, David Brady, Qionghai Dai, and Lu Fang. 2020. PANDA: A Gigapixel-Level Human-Centric Video Dataset. In 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 3265--3275."},{"key":"e_1_3_2_1_41_1","volume-title":"Learning Actor Relation Graphs for Group Activity Recognition. In 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 9956--9966","author":"Wu Jianchao","year":"2019","unstructured":"Jianchao Wu, Limin Wang, Li Wang, Jie Guo, and Gangshan Wu. 2019. Learning Actor Relation Graphs for Group Activity Recognition. In 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 9956--9966."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_42_1","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"e_1_3_2_1_43_1","volume-title":"An Actor-centric Causality Graph for Asynchronous Temporal Inference in Group Activity. In 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 6652--6661","author":"Xie Zhao","year":"2023","unstructured":"Zhao Xie, Tian Gao, Kewei Wu, and Jiao Chang. 2023. An Actor-centric Causality Graph for Asynchronous Temporal Inference in Group Activity. In 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 6652--6661."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_44_1","DOI":"10.1145\/3240508.3240572"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_45_1","DOI":"10.1109\/TPAMI.2020.3034233"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_46_1","DOI":"10.1109\/TPAMI.2018.2874455"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_47_1","DOI":"10.1609\/aaai.v35i4.16437"},{"key":"e_1_3_2_1_48_1","volume-title":"Open-Vocabulary Object Detection Using Captions. In 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 14388--14397","author":"Zareian Alireza","year":"2021","unstructured":"Alireza Zareian, Kevin Dela Rosa, Derek Hao Hu, and Shih-Fu Chang. 2021. Open-Vocabulary Object Detection Using Captions. In 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 14388--14397."},{"key":"e_1_3_2_1_49_1","volume-title":"Self-tuning spectral clustering. Advances in neural information processing systems","author":"Zelnik-Manor Lihi","year":"2004","unstructured":"Lihi Zelnik-Manor and Pietro Perona. 2004. Self-tuning spectral clustering. Advances in neural information processing systems, Vol. 17 (2004)."},{"key":"e_1_3_2_1_50_1","volume-title":"International Conference on Machine Learning (ICML).","author":"Zeng Yan","year":"2022","unstructured":"Yan Zeng, Xinsong Zhang, and Hang Li. 2022. Multi-grained vision language pre-training: Aligning texts with visual concepts. In International Conference on Machine Learning (ICML)."}],"event":{"sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"acronym":"MM '24","name":"MM '24: The 32nd ACM International Conference on Multimedia","location":"Melbourne VIC Australia"},"container-title":["Proceedings of the 32nd ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3664647.3681329","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3664647.3681329","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:17:43Z","timestamp":1750295863000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3664647.3681329"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,28]]},"references-count":50,"alternative-id":["10.1145\/3664647.3681329","10.1145\/3664647"],"URL":"https:\/\/doi.org\/10.1145\/3664647.3681329","relation":{},"subject":[],"published":{"date-parts":[[2024,10,28]]},"assertion":[{"value":"2024-10-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}