{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:53:33Z","timestamp":1782312813701,"version":"3.54.5"},"reference-count":87,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T00:00:00Z","timestamp":1782259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"State Key Laboratory for Novel Software Technology, Nanjing University","award":["KFKT2025A22"],"award-info":[{"award-number":["KFKT2025A22"]}]},{"name":"Natural Science Foundation of Zhejiang Province of China","award":["LMS26F020036 and LY24F020012"],"award-info":[{"award-number":["LMS26F020036 and LY24F020012"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62476238"],"award-info":[{"award-number":["62476238"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Hangzhou Key Scientific Research Program","award":["2025SZD1B10"],"award-info":[{"award-number":["2025SZD1B10"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>Recent advances in action segmentation have greatly enhanced our understanding of complex and dynamic scenes in video content. Despite these improvements, the field continues to face persistent challenges, particularly in terms of model efficiency and the substantial cost associated with manual annotation. In this work, we introduce a novel framework that integrates active learning within hyperbolic space to effectively address these issues. By leveraging the hierarchical representational capacity of hyperbolic space, which is naturally suited for modeling structured data, and combining it with the selective efficiency of active learning, our method introduces hyperbolic uncertainty metrics to guide the targeted selection of the most informative video frames and sequences for annotation. This enables the model to prioritize annotation efforts where they are most impactful. Furthermore, the model iteratively refines pseudo labels using all available annotations, significantly reducing the need for exhaustive labeling while preserving high segmentation accuracy. To further mitigate reliance on precise annotations, we enhance the MS-TCN model by incorporating soft pseudo labels and a weighting mechanism that dynamically adjusts learning based on label confidence, allowing for more robust training in the presence of noisy or weakly labeled data. Extensive experiments conducted on two widely used action segmentation benchmark datasets validate the effectiveness of our approach, demonstrating that it can substantially reduce annotation effort while maintaining overall segmentation performance.<\/jats:p>","DOI":"10.1145\/3797032","type":"journal-article","created":{"date-parts":[[2026,3,30]],"date-time":"2026-03-30T15:03:33Z","timestamp":1774883013000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Hyperbolic Active Learning for Label-Efficient Action Segmentation"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-7613-7628","authenticated-orcid":false,"given":"Jingqiao","family":"Xiu","sequence":"first","affiliation":[{"name":"Nanjing University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8106-9768","authenticated-orcid":false,"given":"Wei","family":"Ji","sequence":"additional","affiliation":[{"name":"Nanjing University, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2510-5282","authenticated-orcid":false,"given":"Menglin","family":"Yang","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-5047-9852","authenticated-orcid":false,"given":"Yufeng","family":"Chen","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8906-4534","authenticated-orcid":false,"given":"Hanbin","family":"Zhao","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7685-5705","authenticated-orcid":false,"given":"Fangfang","family":"Wang","sequence":"additional","affiliation":[{"name":"Hangzhou Normal University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7410-2590","authenticated-orcid":false,"given":"Roger","family":"Zimmermann","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,24]]},"reference":[{"key":"e_1_3_1_2_2","volume-title":"International Conference on Learning Representations","author":"Ash Jordan T.","year":"2020","unstructured":"Jordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, and Alekh Agarwal. 2020. Deep batch active learning by diverse, uncertain gradient lower bounds. In International Conference on Learning Representations."},{"key":"e_1_3_1_3_2","first-page":"4453","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ghadimi Atigh Mina","year":"2022","unstructured":"Mina Ghadimi Atigh, Julian Schoep, Erman Acar, Nanne Van Noord, and Pascal Mettes. 2022. Hyperbolic image segmentation. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 4453\u20134462."},{"key":"e_1_3_1_4_2","first-page":"10351","volume-title":"IEEE\/CVF International Conference on Computer Vision","author":"Bahrami Emad","year":"2023","unstructured":"Emad Bahrami, Gianpiero Francesca, and Juergen Gall. 2023. How much temporal long-term context is needed for action segmentation? In IEEE\/CVF International Conference on Computer Vision, 10351\u201310361."},{"key":"e_1_3_1_5_2","first-page":"12316","volume-title":"Advances in Neural Information Processing Systems","author":"Bai Yushi","year":"2021","unstructured":"Yushi Bai, Zhitao Ying, Hongyu Ren, and Jure Leskovec. 2021. Modeling heterogeneous hierarchies with relation-specific hyperbolic cones. In Advances in Neural Information Processing Systems, 12316\u201312327."},{"key":"e_1_3_1_6_2","volume-title":"Advances in Neural Information Processing Systems","author":"Balazevic Ivana","year":"2019","unstructured":"Ivana Balazevic, Carl Allen, and Timothy Hospedales. 2019. Multi-relational poincar\u00e9 graph embeddings. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19833-5_4"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00976"},{"key":"e_1_3_1_9_2","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1017\/9781009701853.003","volume-title":"Flavors of Geometry","author":"Cannon James W.","year":"1997","unstructured":"James W. Cannon, William J. Floyd, Richard Kenyon, Walter R. Parry. 1997. Hyperbolic geometry. In S. Levy, Flavors of Geometry, Cambridge University Press, 59\u2013116."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_1_11_2","first-page":"4868","volume-title":"Advances in Neural Information Processing Systems","author":"Chami Ines","year":"2019","unstructured":"Ines Chami, Zhitao Ying, Christopher R\u00e9, and Jure Leskovec. 2019. Hyperbolic graph convolutional neural networks. In Advances in Neural Information Processing Systems, 4868\u20134879."},{"key":"e_1_3_1_12_2","first-page":"3546","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chang Chien-Yi","year":"2019","unstructured":"Chien-Yi Chang, De-An Huang, Yanan Sui, Li Fei-Fei, and Juan Carlos Niebles. 2019. D3tw: Discriminative differentiable dynamic time warping for weakly supervised action alignment and segmentation. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3546\u20133555."},{"key":"e_1_3_1_13_2","first-page":"8395","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chang Xiaobin","year":"2021","unstructured":"Xiaobin Chang, Frederick Tung, and Greg Mori. 2021. Learning discriminative prototypes with dynamic time warping. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 8395\u20138404."},{"key":"e_1_3_1_14_2","volume-title":"IEEE Transactions on Intelligent Transportation Systems","author":"Chen Bike","year":"2023","unstructured":"Bike Chen, Wei Peng, Xiaofeng Cao, and Juha R\u00f6ning. 2023. Hyperbolic uncertainty aware semantic segmentation. IEEE Transactions on Intelligent Transportation Systems."},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","first-page":"6","DOI":"10.1201\/9781003214892","volume-title":"International Joint Conference on Artificial Intelligence","author":"Chen Lei","year":"2022","unstructured":"Lei Chen, Muheng Li, Yueqi Duan, Jie Zhou, and Jiwen Lu. 2022. Uncertainty-aware representation learning for action segmentation. In International Joint Conference on Artificial Intelligence, 6."},{"key":"e_1_3_1_16_2","first-page":"11933","volume-title":"Advances in Neural Information Processing Systems","author":"Citovsky Gui","year":"2021","unstructured":"Gui Citovsky, Giulia DeSalvo, Claudio Gentile, Lazaros Karydas, Anand Rajagopalan, Afshin Rostamizadeh, and Sanjiv Kumar. 2021. Batch active learning at scale. In Advances in Neural Information Processing Systems, 11933\u201311944."},{"key":"e_1_3_1_17_2","first-page":"745","volume-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence","author":"Collins Robert T.","year":"2000","unstructured":"Robert T. Collins, Alan J. Lipton, and Takeo Kanade. 2000. Introduction to the special section on video surveillance. IEEE Transactions on Pattern Analysis and Machine Intelligence, 745\u2013746."},{"key":"e_1_3_1_18_2","first-page":"7694","volume-title":"International Conference on Machine Learning","author":"Desai Karan","year":"2023","unstructured":"Karan Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson, and Shanmukha Ramakrishna Vedantam. 2023. Hyperbolic image-text representations. In International Conference on Machine Learning, 7694\u20137731."},{"key":"e_1_3_1_19_2","first-page":"6508","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ding Li","year":"2018","unstructured":"Li Ding and Chenliang Xu. 2018. Weakly-supervised action segmentation with iterative soft boundary assignment. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6508\u20136516."},{"key":"e_1_3_1_20_2","first-page":"7409","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ermolov Aleksandr","year":"2022","unstructured":"Aleksandr Ermolov, Leyla Mirvakhabova, Valentin Khrulkov, Nicu Sebe, and Ivan Oseledets. 2022. Hyperbolic vision transformers: Combining improvements in metric learning. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 7409\u20137419."},{"key":"e_1_3_1_21_2","first-page":"3575","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Abu Farha Yazan","year":"2019","unstructured":"Yazan Abu Farha and Jurgen Gall. 2019. MS-TCN: Multi-stage temporal convolutional network for action segmentation. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 3575\u20133584."},{"key":"e_1_3_1_22_2","first-page":"501","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Fayyaz Mohsen","year":"2020","unstructured":"Mohsen Fayyaz and Jurgen Gall. 2020. Sct: Set constrained temporal transformer for set supervised action segmentation. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 501\u2013510."},{"key":"e_1_3_1_23_2","first-page":"1050","volume-title":"International Conference on Machine Learning","author":"Gal Yarin","year":"2016","unstructured":"Yarin Gal and Zoubin Ghahramani. 2016. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. In International Conference on Machine Learning, 1050\u20131059."},{"key":"e_1_3_1_24_2","first-page":"1183","volume-title":"International Conference on Machine Learning","author":"Gal Yarin","year":"2017","unstructured":"Yarin Gal, Riashat Islam, and Zoubin Ghahramani. 2017. Deep Bayesian active learning with image data. In International Conference on Machine Learning, 1183\u20131192."},{"key":"e_1_3_1_25_2","first-page":"1646","volume-title":"International Conference on Machine Learning","author":"Ganea Octavian","year":"2018","unstructured":"Octavian Ganea, Gary B\u00e9cigneul, and Thomas Hofmann. 2018. Hyperbolic entailment cones for learning hierarchical embeddings. In International Conference on Machine Learning, 1646\u20131655."},{"key":"e_1_3_1_26_2","volume-title":"Advances in Neural Information Processing Systems","author":"Ganea Octavian","year":"2018","unstructured":"Octavian Ganea, Gary B\u00e9cigneul, and Thomas Hofmann. 2018. Hyperbolic neural networks. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_1_27_2","first-page":"16805","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Gao Shang-Hua","year":"2021","unstructured":"Shang-Hua Gao, Qi Han, Zhong-Yu Li, Pai Peng, Liang Wang, and Ming-Ming Cheng. 2021. Global2local: Efficient structure search for video action segmentation. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 16805\u201316814."},{"key":"e_1_3_1_28_2","first-page":"6840","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ge Songwei","year":"2023","unstructured":"Songwei Ge, Shlok Mishra, Simon Kornblith, Chun-Liang Li, and David Jacobs. 2023. Hyperbolic contrastive learning for visual representations beyond objects. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6840\u20136849."},{"key":"e_1_3_1_29_2","first-page":"103","volume-title":"Advances in Neural Information Processing Systems","author":"Ghadimi Atigh Mina","year":"2021","unstructured":"Mina Ghadimi Atigh, Martin Keller-Ressel, and Pascal Mettes. 2021. Hyperbolic Busemann learning with ideal prototypes. In Advances in Neural Information Processing Systems, 103\u2013115."},{"key":"e_1_3_1_30_2","first-page":"11","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Guo Yunhui","year":"2022","unstructured":"Yunhui Guo, Xudong Wang, Yubei Chen, and Stella X. Yu. 2022. Clipped hyperbolic classifiers are super-hyperbolic classifiers. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 11\u201320."},{"key":"e_1_3_1_31_2","first-page":"5112","volume-title":"Advances in Neural Information Processing Systems","author":"Hsu Joy","year":"2021","unstructured":"Joy Hsu, Jeffrey Gu, Gong Wu, Wah Chiu, and Serena Yeung. 2021. Capturing implicit hierarchical structure in 3D biomedical images with self-supervised hyperbolic representations. Advances in Neural Information Processing Systems, 5112\u20135123."},{"key":"e_1_3_1_32_2","first-page":"2322","volume-title":"IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Ishikawa Yuchi","year":"2021","unstructured":"Yuchi Ishikawa, Seito Kasai, Yoshimitsu Aoki, and Hirokatsu Kataoka. 2021. Alleviating over-segmentation errors by detecting action boundaries. In IEEE\/CVF Winter Conference on Applications of Computer Vision, 2322\u20132331."},{"key":"e_1_3_1_33_2","unstructured":"Will Kay Joao Carreira Karen Simonyan Brian Zhang Chloe Hillier Sudheendra Vijayanarasimhan Fabio Viola Tim Green Trevor Back Paul Natsev et al. 2017. The kinetics human action video dataset. arXiv:1705.06950. Retrieved from https:\/\/arxiv.org\/abs\/1705.06950"},{"key":"e_1_3_1_34_2","first-page":"10619","volume-title":"IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Khan Hamza","year":"2022","unstructured":"Hamza Khan, Sanjay Haresh, Awais Ahmed, Shakeeb Siddiqui, Andrey Konin, M. Zeeshan Zia, and Quoc-Huy Tran. 2022. Timestamp-supervised action segmentation with graph convolutional networks. In IEEE\/RSJ International Conference on Intelligent Robots and Systems, 10619\u201310626."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00645"},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","first-page":"036106","DOI":"10.1103\/PhysRevE.82.036106","article-title":"Hyperbolic geometry of complex networks","author":"Krioukov Dmitri","year":"2010","unstructured":"Dmitri Krioukov, Fragkiskos Papadopoulos, Maksim Kitsak, Amin Vahdat, and Mari\u00e1n Bogun\u00e1. 2010. Hyperbolic geometry of complex networks. Physical Review. E, Statistical, Nonlinear, and Soft Matter Physics 82 (2010), 036106.","journal-title":"Physical Review. E, Statistical, Nonlinear, and Soft Matter Physics"},{"key":"e_1_3_1_37_2","first-page":"780","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Kuehne Hilde","year":"2014","unstructured":"Hilde Kuehne, Ali Arslan, and Thomas Serre. 2014. The language of actions: Recovering the syntax and semantics of goal-directed human activities. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 780\u2013787."},{"key":"e_1_3_1_38_2","first-page":"156","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lea Colin","year":"2017","unstructured":"Colin Lea, Michael D. Flynn, Rene Vidal, Austin Reiter, and Gregory D. Hager. 2017. Temporal convolutional networks for action segmentation and detection. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 156\u2013165."},{"key":"e_1_3_1_39_2","first-page":"6243","volume-title":"IEEE\/CVF International Conference on Computer Vision","author":"Li Jun","year":"2019","unstructured":"Jun Li, Peng Lei, and Sinisa Todorovic. 2019. Weakly supervised energy-based learning for action segmentation. In IEEE\/CVF International Conference on Computer Vision, 6243\u20136251."},{"key":"e_1_3_1_40_2","first-page":"10820","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Jun","year":"2020","unstructured":"Jun Li and Sinisa Todorovic. 2020. Set-constrained Viterbi for set-supervised action segmentation. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 10820\u201310829."},{"key":"e_1_3_1_41_2","first-page":"9806","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Jun","year":"2021","unstructured":"Jun Li and Sinisa Todorovic. 2021. Anchor-constrained Viterbi for set-supervised action segmentation. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 9806\u20139815."},{"key":"e_1_3_1_42_2","volume-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence","author":"Li Shi-Jie","year":"2020","unstructured":"Shi-Jie Li, Yazan AbuFarha, Yun Liu, Ming-Ming Cheng, and Juergen Gall. 2020. MS-TCN++: Multi-stage temporal convolutional network for action segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence."},{"key":"e_1_3_1_43_2","first-page":"8365","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Zhe","year":"2021","unstructured":"Zhe Li, Yazan Abu Farha, and Jurgen Gall. 2021. Temporal action segmentation from timestamp supervision. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 8365\u20138374."},{"key":"e_1_3_1_44_2","first-page":"10139","volume-title":"IEEE\/CVF International Conference on Computer Vision","author":"Liu Daochang","year":"2023","unstructured":"Daochang Liu, Qiyue Li, Anh-Dung Dinh, Tingting Jiang, Mubarak Shah, and Chang Xu. 2023. Diffusion action segmentation. In IEEE\/CVF International Conference on Computer Vision, 10139\u201310149."},{"key":"e_1_3_1_45_2","first-page":"6503","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu Kaiyuan","year":"2023","unstructured":"Kaiyuan Liu, Yunheng Li, Shenglan Liu, Chenwei Tan, and Zihang Shao. 2023. Reducing the label bias for timestamp supervised temporal action segmentation. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6503\u20136513."},{"key":"e_1_3_1_46_2","first-page":"8230","volume-title":"Advances in Neural Information Processing Systems","author":"Liu Qi","year":"2019","unstructured":"Qi Liu, Maximilian Nickel, and Douwe Kiela. 2019. Hyperbolic graph neural networks. In Advances in Neural Information Processing Systems, 8230\u20138241."},{"key":"e_1_3_1_47_2","first-page":"9273","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu Shaoteng","year":"2020","unstructured":"Shaoteng Liu, Jingjing Chen, Liangming Pan, Chong-Wah Ngo, Tat-Seng Chua, and Yu-Gang Jiang. 2020. Hyperbolic visual embedding learning for zero-shot recognition. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 9273\u20139281."},{"key":"e_1_3_1_48_2","first-page":"1141","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Long Teng","year":"2020","unstructured":"Teng Long, Pascal Mettes, Heng Tao Shen, and Cees G. M. Snoek. 2020. Searching for actions on the hyperbole. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 1141\u20131150."},{"key":"e_1_3_1_49_2","first-page":"8085","volume-title":"IEEE\/CVF International Conference on Computer Vision","author":"Lu Zijia","year":"2021","unstructured":"Zijia Lu and Ehsan Elhamifar. 2021. Weakly-supervised action segmentation and alignment via transcript-aware union-of-subspaces learning. In IEEE\/CVF International Conference on Computer Vision, 8085\u20138095."},{"key":"e_1_3_1_50_2","first-page":"19903","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lu Zijia","year":"2022","unstructured":"Zijia Lu and Ehsan Elhamifar. 2022. Set-supervised action learning in procedural task videos via pairwise order consistency. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 19903\u201319913."},{"key":"e_1_3_1_51_2","doi-asserted-by":"crossref","first-page":"3484","DOI":"10.1007\/s11263-024-02043-5","article-title":"Hyperbolic deep learning in computer vision: A survey","author":"Mettes Pascal","year":"2024","unstructured":"Pascal Mettes, Mina Ghadimi Atigh, Martin Keller-Ressel, Jeffrey Gu, and Serena Yeung. 2024. Hyperbolic deep learning in computer vision: A survey. International Journal of Computer Vision 132 (2024), 3484\u20133508.","journal-title":"International Journal of Computer Vision"},{"key":"e_1_3_1_52_2","first-page":"9915","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Moltisanti Davide","year":"2019","unstructured":"Davide Moltisanti, Sanja Fidler, and Dima Damen. 2019. Action recognition from single timestamp supervision in untrimmed videos. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 9915\u20139924."},{"key":"e_1_3_1_53_2","first-page":"33741","volume-title":"Advances in Neural Information Processing Systems","author":"Montanaro Antonio","year":"2022","unstructured":"Antonio Montanaro, Diego Valsesia, and Enrico Magli. 2022. Rethinking the compositionality of point clouds through regularization in the hyperbolic space. In Advances in Neural Information Processing Systems, 33741\u201333753."},{"key":"e_1_3_1_54_2","volume-title":"Advances in Neural Information Processing Systems","author":"Nickel Maximillian","year":"2017","unstructured":"Maximillian Nickel and Douwe Kiela. 2017. Poincar\u00e9 embeddings for learning hierarchical representations. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_1_55_2","first-page":"3779","volume-title":"International Conference on Machine Learning","author":"Nickel Maximillian","year":"2018","unstructured":"Maximillian Nickel and Douwe Kiela. 2018. Learning continuous hierarchies in the Lorentz model of hyperbolic geometry. In International Conference on Machine Learning, 3779\u20133788."},{"key":"e_1_3_1_56_2","doi-asserted-by":"crossref","first-page":"108764","DOI":"10.1016\/j.patcog.2022.108764","article-title":"Maximization and restoration: Action segmentation through dilation passing and temporal reconstruction","author":"Park Junyong","year":"2022","unstructured":"Junyong Park, Daekyum Kim, Sejoon Huh, and Sungho Jo. 2022. Maximization and restoration: Action segmentation through dilation passing and temporal reconstruction. Pattern Recognition (2022), 108764.","journal-title":"Pattern Recognition"},{"key":"e_1_3_1_57_2","first-page":"10023","volume-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence","author":"Peng Wei","year":"2021","unstructured":"Wei Peng, Tuomas Varanka, Abdelrahman Mostafa, Henglin Shi, and Guoying Zhao. 2021. Hyperbolic deep neural networks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 10023\u201310044."},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19772-7_17"},{"key":"e_1_3_1_59_2","first-page":"754","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Richard Alexander","year":"2017","unstructured":"Alexander Richard, Hilde Kuehne, and Juergen Gall. 2017. Weakly supervised action learning with rnn based fine-to-coarse modeling. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 754\u2013763."},{"key":"e_1_3_1_60_2","first-page":"5987","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Richard Alexander","year":"2018","unstructured":"Alexander Richard, Hilde Kuehne, and Juergen Gall. 2018. Action sets: Weakly supervised action segmentation without ordering constraints. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 5987\u20135996."},{"key":"e_1_3_1_61_2","first-page":"7386","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Richard Alexander","year":"2018","unstructured":"Alexander Richard, Hilde Kuehne, Ahsan Iqbal, and Juergen Gall. 2018. Neuralnetwork-viterbi: A framework for weakly supervised video learning. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 7386\u20137395."},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0851-8"},{"key":"e_1_3_1_63_2","volume-title":"International Conference on Learning Representations","author":"Shimizu Ryohei","year":"2021","unstructured":"Ryohei Shimizu, Yusuke Mukuta, and Tatsuya Harada. 2021. Hyperbolic neural networks++. In International Conference on Learning Representations."},{"key":"e_1_3_1_64_2","first-page":"1308","volume-title":"International Conference on Artificial Intelligence and Statistics","author":"Shui Changjian","year":"2020","unstructured":"Changjian Shui, Fan Zhou, Christian Gagn\u00e9, and Boyu Wang. 2020. Deep active learning: Unified and principled method for query and training. In International Conference on Artificial Intelligence and Statistics, 1308\u20131318."},{"key":"e_1_3_1_65_2","doi-asserted-by":"crossref","first-page":"11484","DOI":"10.1109\/TPAMI.2023.3284080","article-title":"C2F-TCN: A framework for semi-and fully-supervised temporal action segmentation","author":"Singhania Dipika","year":"2023","unstructured":"Dipika Singhania, Rahul Rahaman, and Angela Yao. 2023. C2F-TCN: A framework for semi-and fully-supervised temporal action segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (2023), 11484\u201311501.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_66_2","volume-title":"British Machine Vision Conference","author":"Souri Yaser","year":"2022","unstructured":"Yaser Souri, Yazan Abu Farha, Emad Bahrami, Gianpiero Francesca, and Juergen Gall. 2022. Robust action segmentation from timestamp supervision. In British Machine Vision Conference."},{"key":"e_1_3_1_67_2","doi-asserted-by":"crossref","first-page":"282","DOI":"10.1007\/978-3-030-92659-5_18","volume-title":"German Conference on Pattern Recognition","author":"Souri Yaser","year":"2021","unstructured":"Yaser Souri, Yazan Abu Farha, Fabien Despinoy, Gianpiero Francesca, and Juergen Gall. 2021. Fifa: Fast inference approximation for action segmentation. In German Conference on Pattern Recognition, 282\u2013296."},{"key":"e_1_3_1_68_2","first-page":"6196","volume-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence","author":"Souri Yaser","year":"2021","unstructured":"Yaser Souri, Mohsen Fayyaz, Luca Minciullo, Gianpiero Francesca, and Juergen Gall. 2021. Fast weakly supervised action segmentation using mutual consistency. In IEEE Transactions on Pattern Analysis and Machine Intelligence, 6196\u20136208."},{"key":"e_1_3_1_69_2","first-page":"161","volume-title":"European Conference on Computer Vision","author":"Su Yuhao","year":"2024","unstructured":"Yuhao Su and Ehsan Elhamifar. 2024. Two-stage active learning for efficient temporal action segmentation. In European Conference on Computer Vision. Springer, 161\u2013183."},{"key":"e_1_3_1_70_2","first-page":"882","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Vats Kanav","year":"2020","unstructured":"Kanav Vats, Mehrnaz Fani, Pascale Walters, David A. Clausi, and John Zelek. 2020. Event detection in coarsely annotated sports videos via parallel multi-receptive field 1D convolutions. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, 882\u2013883."},{"key":"e_1_3_1_71_2","first-page":"983","volume-title":"The Visual Computer","volume":"29","author":"Vishwakarma Sarvesh","year":"2013","unstructured":"Sarvesh Vishwakarma and Anupam Agrawal. 2013. A survey on activity recognition and behavior understanding in video surveillance. The Visual Computer 29 (2013), 983\u20131009."},{"key":"e_1_3_1_72_2","first-page":"8566","volume-title":"AAAI Conference on Artificial Intelligence","author":"Wang Tianyang","year":"2022","unstructured":"Tianyang Wang, Xingjian Li, Pengkun Yang, Guosheng Hu, Xiangrui Zeng, Siyu Huang, Cheng-Zhong Xu, and Min Xu. 2022. Boosting active learning via improving test performance. In AAAI Conference on Artificial Intelligence, 8566\u20138574."},{"key":"e_1_3_1_73_2","first-page":"2603","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Weng Zhenzhen","year":"2021","unstructured":"Zhenzhen Weng, Mehmet Giray Ogut, Shai Limonchik, and Serena Yeung. 2021. Unsupervised discovery of the long-tail in instance segmentation using hierarchical self-supervision. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2603\u20132612."},{"key":"e_1_3_1_74_2","volume-title":"International Conference on Learning Representations","author":"Xiu Jingqiao","year":"2026","unstructured":"Jingqiao Xiu, Fangzhou Hong, Yicong Li, Mengze Li, Wentao Wang, Sirui Han, Liang Pan, and Ziwei Liu. 2026. Egotwin: Dreaming body and view in first person. In International Conference on Learning Representations."},{"key":"e_1_3_1_75_2","first-page":"9271","volume-title":"ACM International Conference on Multimedia","author":"Xiu Jingqiao","year":"2024","unstructured":"Jingqiao Xiu, Mengze Li, Wei Ji, Jingyuan Chen, Hanbin Zhao, Shin\u2019ichi Satoh, and Roger Zimmermann. 2024. Hierarchical debiasing and noisy correction for cross-domain video tube retrieval. In ACM International Conference on Multimedia, 9271\u20139280."},{"key":"e_1_3_1_76_2","first-page":"8788","volume-title":"AAAI Conference on Artificial Intelligence","author":"Xiu Jingqiao","year":"2025","unstructured":"Jingqiao Xiu, Mengze Li, Zongxin Yang, Wei Ji, Yifang Yin, and Roger Zimmermann. 2025. Few-shot incremental learning via foreground aggregation and knowledge transfer for audio-visual semantic segmentation. In AAAI Conference on Artificial Intelligence, 8788\u20138796."},{"key":"e_1_3_1_77_2","first-page":"27435","volume-title":"IEEE\/CVF International Conference on Computer Vision","author":"Xiu Jingqiao","year":"2025","unstructured":"Jingqiao Xiu, Yicong Li, Na Zhao, Han Fang, Xiang Wang, and Angela Yao. 2025. Geometric alignment and prior modulation for view-guided point cloud completion on unseen categories. In IEEE\/CVF International Conference on Computer Vision, 27435\u201327444."},{"key":"e_1_3_1_78_2","first-page":"14890","volume-title":"Advances in Neural Information Processing Systems","author":"Xu Ziwei","year":"2022","unstructured":"Ziwei Xu, Yogesh Rawat, Yongkang Wong, Mohan S. Kankanhalli, and Mubarak Shah. 2022. Don\u2019t pour cereal into coffee: Differentiable temporal logic for temporal action segmentation. In Advances in Neural Information Processing Systems, 14890\u201314903."},{"key":"e_1_3_1_79_2","first-page":"2212","volume-title":"ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Yang Menglin","year":"2022","unstructured":"Menglin Yang, Zhihao Li, Min Zhou, Jiahong Liu, and Irwin King. 2022. HICF: Hyperbolic informative collaborative filtering. In ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2212\u20132221."},{"key":"e_1_3_1_80_2","first-page":"1975","volume-title":"ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Yang Menglin","year":"2021","unstructured":"Menglin Yang, Min Zhou, Marcus Kalander, Zengfeng Huang, and Irwin King. 2021. Discrete-time temporal network embedding via implicit hierarchical learning in hyperbolic space. In ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 1975\u20131985."},{"key":"e_1_3_1_81_2","unstructured":"Menglin Yang Min Zhou Zhihao Li Jiahong Liu Lujia Pan Hui Xiong and Irwin King. 2022. Hyperbolic graph neural networks: A review of methods and applications. arXiv:2202.13852. Retrieved from https:\/\/arxiv.org\/abs\/1705.06950"},{"key":"e_1_3_1_82_2","volume-title":"ACM Web Conference","author":"Yang Menglin","year":"2022","unstructured":"Menglin Yang, Min Zhou, Jiahong Liu, Defu Lian, and Irwin King. 2022. HRCF: Enhancing collaborative filtering via hyperbolic geometric regularization. In ACM Web Conference."},{"key":"e_1_3_1_83_2","doi-asserted-by":"publisher","DOI":"10.5244\/C.35.49"},{"key":"e_1_3_1_84_2","first-page":"575","volume-title":"IEEE International Conference on Data Mining","author":"Yin Changchang","year":"2017","unstructured":"Changchang Yin, Buyue Qian, Shilei Cao, Xiaoyu Li, Jishang Wei, Qinghua Zheng, and Ian Davidson. 2017. Deep similarity-based batch mode active learning with exploration-exploitation. In IEEE International Conference on Data Mining, 575\u2013584."},{"key":"e_1_3_1_85_2","first-page":"93","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yoo Donggeun","year":"2019","unstructured":"Donggeun Yoo and In So Kweon. 2019. Learning loss for active learning. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 93\u2013102."},{"key":"e_1_3_1_86_2","first-page":"3723","volume-title":"International Joint Conference on Artificial Intelligence","author":"Zhang Baoquan","year":"2022","unstructured":"Baoquan Zhang, Hao Jiang, Shanshan Feng, Xutao Li, Yunming Ye, and Rui Ye. 2022. Hyperbolic knowledge transfer with class hierarchy for few-shot learning. International Joint Conference on Artificial Intelligence, 3723\u20133729."},{"key":"e_1_3_1_87_2","unstructured":"Fedor Zhdanov. 2019. Diverse mini-batch active learning. arXiv:1901.05954. Retrieved from https:\/\/arxiv.org\/abs\/1705.06950"},{"key":"e_1_3_1_88_2","volume-title":"Findings of Empirical Methods in Natural Language Processing","author":"Zhu Yudong","year":"2020","unstructured":"Yudong Zhu, Di Zhou, Jinghui Xiao, Xin Jiang, Xiao Chen, and Qun Liu. 2020. Hypertext: Endowing fasttext with hyperbolic geometry. In Findings of Empirical Methods in Natural Language Processing."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3797032","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T14:41:04Z","timestamp":1782312064000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3797032"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,24]]},"references-count":87,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3797032"],"URL":"https:\/\/doi.org\/10.1145\/3797032","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,24]]},"assertion":[{"value":"2025-04-09","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}