{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T17:45:26Z","timestamp":1776879926405,"version":"3.51.2"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2025,4,9]],"date-time":"2025-04-09T00:00:00Z","timestamp":1744156800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2024YFB3311602"],"award-info":[{"award-number":["2024YFB3311602"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272144"],"award-info":[{"award-number":["62272144"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100003995","name":"Anhui Provincial Natural Science Foundation","doi-asserted-by":"crossref","award":["2408085J040"],"award-info":[{"award-number":["2408085J040"]}],"id":[{"id":"10.13039\/501100003995","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Anhui Provincial Major Science and Technology","award":["202423k09020001"],"award-info":[{"award-number":["202423k09020001"]}]},{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities","doi-asserted-by":"crossref","award":["JZ2024HGTG0309, JZ2024AHST0337"],"award-info":[{"award-number":["JZ2024HGTG0309, JZ2024AHST0337"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,4,30]]},"abstract":"<jats:p>\n            Repetitive Action Counting (RAC) is a critical and challenging task in video analysis, aiming to count the number of repeated actions in videos accurately. Existing methods typically generate a Temporal Self-similarity Matrix (TSM) as an intermediate representation to predict the number of repetitive actions. While this simplifies the process, it often overlooks the variable lengths between action cycles and the phenomenon of motion interruptions. The period inconsistency problem caused by the change in the action period and the motion interruption problem resulting from the motion pause are the two main challenges that affect the accuracy of RAC in complex scenes. To address these challenges, we propose a novel framework. First, we construct a boundary-aware encoder equipped with a temporal pyramid structure to build multi-scale video features, capturing the period information of different lengths of repetitive actions to solve the period inconsistency problem. Next, a cycle and boundary attention module is followed by each layer in the pyramid to enhance these multi-scale features with periodic and event boundary information. Finally, we design a gated density estimator to generate the actionness score for each frame that reflects the probability of the corresponding time point being within the motion cycle. These scores are used to weight features to reduce the impact of noise frames without actions present and solve the motion interruption problem for better density prediction. Extensive experiments conducted on public datasets demonstrate the effectiveness of our method. The source code will be available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/zqzhang2023\/TBANRAC\">https:\/\/github.com\/zqzhang2023\/TBANRAC<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3712602","type":"journal-article","created":{"date-parts":[[2025,1,21]],"date-time":"2025-01-21T12:09:22Z","timestamp":1737461362000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Temporal Boundary Awareness Network for Repetitive Action Counting"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-8487-9127","authenticated-orcid":false,"given":"Zhenqiang","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Computer Science and Information Engineering, School of Artificial Intelligence, Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5083-2145","authenticated-orcid":false,"given":"Kun","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Engineering, School of Artificial Intelligence, Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6313-2543","authenticated-orcid":false,"given":"Shengeng","family":"Tang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Engineering, School of Artificial Intelligence, Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8818-6740","authenticated-orcid":false,"given":"Yanyan","family":"Wei","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Engineering, School of Artificial Intelligence, Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1142-6434","authenticated-orcid":false,"given":"Fei","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Engineering, School of Artificial Intelligence, Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6402-7593","authenticated-orcid":false,"given":"Jinxing","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Engineering, School of Artificial Intelligence, Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2594-254X","authenticated-orcid":false,"given":"Dan","family":"Guo","sequence":"additional","affiliation":[{"name":"School of Computer Science and Information Engineering, School of Artificial Intelligence, Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,4,9]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"1","volume-title":"Proceedings of the 2008 19th International Conference on Pattern Recognition","author":"Azy Ousman","year":"2008","unstructured":"Ousman Azy and Narendra Ahuja. 2008. Segmentation of periodically moving objects. In Proceedings of the 2008 19th International Conference on Pattern Recognition, 1\u20134."},{"key":"e_1_3_1_3_2","doi-asserted-by":"crossref","first-page":"284","DOI":"10.1007\/3-540-45344-X_42","volume-title":"Proceedings of the 3rd International Conference on Audio-and Video-Based Biometric Person Authentication (AVBPA \u201901)","author":"BenAbdelkader Chiraz","year":"2001","unstructured":"Chiraz BenAbdelkader, Ross Cutler, Harsh Nanda, and Larry Davis. 2001. Eigengait: Motion-based recognition of people using image self-similarity. In Proceedings of the 3rd International Conference on Audio-and Video-Based Biometric Person Authentication (AVBPA \u201901), 284\u2013294."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1155\/S1110865704309236"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2007.1042"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2005.864234"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2022.3193752"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3014555"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2959977"},{"key":"e_1_3_1_10_2","first-page":"167","volume-title":"British Machine Vision Conference","volume":"1","author":"Chetverikov Dmitry","year":"2006","unstructured":"Dmitry Chetverikov and S\u00e1ndor Fazekas. 2006. On motion periodicity of dynamic textures. In British Machine Vision Conference, Vol. 1, 167\u2013176."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/34.868681"},{"key":"e_1_3_1_12_2","first-page":"20041","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Dai Rui","year":"2022","unstructured":"Rui Dai, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo, and Fran\u00e7ois Br\u00e9mond. 2022. Ms-tct: Multi-scale temporal convtransformer for action detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 20041\u201320051."},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3643815"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/3567828"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01040"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00028"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3358415"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2011.941846"},{"key":"e_1_3_1_19_2","first-page":"19013","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Hu Huazhang","year":"2022","unstructured":"Huazhang Hu, Sixun Dong, Yiqun Zhao, Dongze Lian, Zhengxin Li, and Shenghua Gao. 2022. Transrac: Encoding multi-scale temporal correlation with transformers for repetitive action counting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 19013\u201319022."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01404"},{"key":"e_1_3_1_21_2","unstructured":"Van Thong Huynh Hyung-Jeong Yang Guee-Sang Lee and Soo-Hyung Kim. 2023. Generic event boundary detection in video with pyramid features. arXiv:2301.04288. Retrieved from https:\/\/arxiv.org\/abs\/2301.04288"},{"issue":"1","key":"e_1_3_1_22_2","doi-asserted-by":"crossref","first-page":"172","DOI":"10.1109\/TPAMI.2010.68","article-title":"View-independent action recognition from temporal self-similarities","volume":"33","author":"Junejo Imran N.","year":"2010","unstructured":"Imran N. Junejo, Emilie Dexter, Ivan Laptev, and Patrick Perez. 2010. View-independent action recognition from temporal self-similarities. IEEE Transactions on Pattern Analysis and Machine Intelligence 33, 1 (2010), 172\u2013185.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_23_2","first-page":"20073","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Kang Hyolim","year":"2022","unstructured":"Hyolim Kang, Jinwoo Kim, Taehyun Kim, and Seon Joo Kim. 2022. Uboco: Unsupervised boundary contrastive learning for generic event boundary detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 20073\u201320082."},{"key":"e_1_3_1_24_2","first-page":"10286","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Kim Jihwan","year":"2023","unstructured":"Jihwan Kim, Miso Lee, and Jae-Pil Heo. 2023. Self-feedback detr for temporal action detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 10286\u201310296."},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","first-page":"163","DOI":"10.1007\/978-3-642-40261-6_19","volume-title":"Proceedings of the 15th International Conference on Computer Analysis of Images and Patterns (CAIP \u201913), Part I","author":"K\u00f6rner Marco","year":"2013","unstructured":"Marco K\u00f6rner and Joachim Denzler. 2013. Temporal self-similarity for appearance-based action recognition in multi-view setups. In Proceedings of the 15th International Conference on Computer Analysis of Images and Patterns (CAIP \u201913), Part I, 163\u2013171."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.346"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1049\/cvi2.12193"},{"key":"e_1_3_1_28_2","unstructured":"Congcong Li Xinyao Wang Dexiang Hong Yufei Wang Libo Zhang Tiejian Luo and Longyin Wen. 2022. Structured context transformer for generic event boundary detection. arXiv:2206.02985. Retrieved from https:\/\/arxiv.org\/abs\/2206.02985"},{"key":"e_1_3_1_29_2","first-page":"563","volume-title":"Proceedings of the International Conference on Intelligent Robotics and Applications","author":"Li Jianing","year":"2023","unstructured":"Jianing Li, Bowen Chen, Zhiyong Wang, and Honghai Liu. 2023. Full resolution repetition counting. In Proceedings of the International Conference on Intelligent Robotics and Applications, 563\u2013574."},{"key":"e_1_3_1_30_2","unstructured":"Kun Li Dan Guo Guoliang Chen Xinge Peng and Meng Wang. 2023. Joint skeletal and semantic embedding loss for micro-gesture classification. arXiv:2307.10624. Retrieved from https:\/\/arxiv.org\/abs\/2307.10624"},{"key":"e_1_3_1_31_2","first-page":"1902","volume-title":"Proceedings of the 35th AAAI Conference on Artificial Intelligence","author":"Li Kun","year":"2021","unstructured":"Kun Li, Dan Guo, and Meng Wang. 2021. Proposal-free video grounding with contextual pyramid network. In Proceedings of the 35th AAAI Conference on Artificial Intelligence, 1902\u20131910."},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-022-3783-3"},{"key":"e_1_3_1_33_2","unstructured":"Xiaoxiao Li Vivek Singh Yifan Wu Klaus Kirchberg James Duncan and Ankur Kapoor. 2018. Repetitive motion estimation network: Recover cardiac and respiratory signal from thoracic imaging. arXiv:1811.03343. Retrieved from https:\/\/arxiv.org\/abs\/1811.03343"},{"key":"e_1_3_1_34_2","first-page":"6499","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Li Xinjie","year":"2024","unstructured":"Xinjie Li and Huijuan Xu. 2024. Repetitive action counting with motion feature learning. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 6499\u20136508."},{"key":"e_1_3_1_35_2","first-page":"3306","volume-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence","author":"Li Zhangbin","year":"2024","unstructured":"Zhangbin Li, Dan Guo, Jinxing Zhou, Jing Zhang, and Meng Wang. 2024. Object-aware adaptive-positivity learning for audio-visual question answering. In Proceedings of the 38th AAAI Conference on Artificial Intelligence, 3306\u20133314."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01345"},{"key":"e_1_3_1_38_2","first-page":"2187","volume-title":"Proceedings of the 2024 IEEE International Conference on Image Processing","author":"Luo Yanan","year":"2024","unstructured":"Yanan Luo, Jinhui Yi, Yazan Abu Farha, Moritz Wolter, and Juergen Gall. 2024. Rethinking temporal self-similarity for repetitive action counting. In Proceedings of the 2024 IEEE International Conference on Image Processing, 2187\u20132193."},{"key":"e_1_3_1_39_2","first-page":"923","volume-title":"Proceedings of the 2018 25th IEEE International Conference on Image Processing","author":"Panagiotakis Costas","year":"2018","unstructured":"Costas Panagiotakis, Giorgos Karvounas, and Antonis Argyros. 2018. Unsupervised detection of periodic segments in videos. In Proceedings of the 2018 25th IEEE International Conference on Image Processing, 923\u2013927."},{"key":"e_1_3_1_40_2","first-page":"1","volume-title":"Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition","author":"Pogalin Erik","year":"2008","unstructured":"Erik Pogalin, Arnold W. M. Smeulders, and Andrew H. C. Thean. 2008. Visual quasi-periodicity. In Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition, 1\u20138."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3680670"},{"key":"e_1_3_1_42_2","unstructured":"Khurram Soomro Amir Roshan Zamir and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv:1212.0402. Retrieved from https:\/\/arxiv.org\/abs\/1212.0402"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.3390\/s19030714"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3654671"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3117124"},{"key":"e_1_3_1_46_2","first-page":"402","volume-title":"Proceedings of the 16th European Conference on Computer Vision (ECCV \u2019 20)","author":"Teed Zachary","year":"2020","unstructured":"Zachary Teed and Jia Deng. 2020. Raft: Recurrent all-pairs field transforms for optical flow. In Proceedings of the 16th European Conference on Computer Vision (ECCV \u2019 20), Part II, 402\u2013419."},{"key":"e_1_3_1_47_2","first-page":"176","volume-title":"Proceedings of the 2005 7th IEEE Workshops on Applications of Computer Vision (WACV\/MOTION \u201905)","volume":"1","author":"Thangali Ashwin","year":"2005","unstructured":"Ashwin Thangali and Stan Sclaroff. 2005. Periodic motion detection and estimation via space-time sampling. In Proceedings of the 2005 7th IEEE Workshops on Applications of Computer Vision (WACV\/MOTION \u201905), Vol. 1, 176\u2013182."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2018.00143"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1016\/0031-3203(94)90079-5"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-023-15196-1"},{"key":"e_1_3_1_51_2","first-page":"5345","volume-title":"Proceedings of the 38th AAAI Conference on Artificial Intelligence","author":"Wang Fei","year":"2024","unstructured":"Fei Wang, Dan Guo, Kun Li, and Meng Wang. 2024. Eulermormer: Robust Eulerian motion magnification via dynamic filtering within transformer. In Proceedings of the 38th AAAI Conference on Artificial Intelligence, 5345\u20135353."},{"key":"e_1_3_1_52_2","first-page":"18984","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Fei","year":"2024","unstructured":"Fei Wang, Dan Guo, Kun Li, Zhun Zhong, and Meng Wang. 2024. Frequency decoupling for motion magnification via multi-level isomorphic architecture. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 18984\u201318994."},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.441"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.compag.2024.109169"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3567827"},{"key":"e_1_3_1_56_2","first-page":"3054","volume-title":"Proceedings of the 36th AAAI Conference on Artificial Intelligence","author":"Yang Haosen","year":"2022","unstructured":"Haosen Yang, Wenhao Wu, Lining Wang, Sheng Jin, Boyang Xia, Hongxun Yao, and Hujie Huang. 2022. Temporal action proposal generation with background constraint. In Proceedings of the 36th AAAI Conference on Artificial Intelligence, 3054\u20133062."},{"key":"e_1_3_1_57_2","unstructured":"Xiangchen Yin Donglin Di Lei Fan Hao Li Chen Wei Xiaofei Gou Yang Song Xiao Sun and Xun Yang. 2024. Grpose: Learning graph relations for human image generation with pose priors. arXiv:2408.16540. Retrieved from https:\/\/arxiv.org\/abs\/2408.16540"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00075"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-023-01921-8"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01385"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1145\/3361845"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3712602","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3712602","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:37Z","timestamp":1750295917000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3712602"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,4,9]]},"references-count":60,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,4,30]]}},"alternative-id":["10.1145\/3712602"],"URL":"https:\/\/doi.org\/10.1145\/3712602","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,4,9]]},"assertion":[{"value":"2024-10-13","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-10","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}