{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,9]],"date-time":"2026-03-09T09:17:12Z","timestamp":1773047832126,"version":"3.50.1"},"reference-count":38,"publisher":"PeerJ","license":[{"start":{"date-parts":[[2026,3,9]],"date-time":"2026-03-09T00:00:00Z","timestamp":1773014400000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Innovation of Policing Science and Technology, Fujian Province","award":["2024Y0061"],"award-info":[{"award-number":["2024Y0061"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62072110 and U21A20472"],"award-info":[{"award-number":["62072110 and U21A20472"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"abstract":"<jats:p>As a recently proposed task on video understanding, action spotting aims to locate and classify action in a long video clip and it can be applied for automatically generating video summaries and highlights. To resolve this problem, this article proposes a novel framework capable of capturing both intra- and inter-video contextual information. In particular, we present a transformer based model that views the task as a set prediction problem that aims to match the set of predicted action instances and the set of ground-truths. It is able to capture long-range intra-video temporal information and discovers causal relationships between actions. Next, based on the observation that actions of the same type recur in different videos, we propose to exploit the inter-video contextual information from dataset. To do so, we design an action memory module which stores the compact feature representation of each action class during training, so as to improve action recognition and localization performance. We evaluate our model on public benchmark and demonstrate that our model outperforms the state-of-the-art methods by a large margin.<\/jats:p>","DOI":"10.7717\/peerj-cs.3667","type":"journal-article","created":{"date-parts":[[2026,3,9]],"date-time":"2026-03-09T08:20:24Z","timestamp":1773044424000},"page":"e3667","source":"Crossref","is-referenced-by-count":0,"title":["Intra- and inter-video context-aware action spotting"],"prefix":"10.7717","volume":"12","author":[{"given":"Yi","family":"Chen","sequence":"first","affiliation":[{"name":"Department of Computer and Information Security Management, Fujian Police College, Fuzhou, Fujian, China"},{"name":"Collaborative Innovation Research Center of Intelligent Policing, Fujian Police College, Fuzhou, Fujian, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-5697-1746","authenticated-orcid":true,"given":"Jiaxin","family":"Cai","sequence":"additional","affiliation":[{"name":"College of Computer and Data Science, Fuzhou University, Fuzhou, Fujian, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinyang","family":"Lin","sequence":"additional","affiliation":[{"name":"Xiamen Zhonglian Century Co., Ltd., Xiamen, Fujian, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiang","family":"Lin","sequence":"additional","affiliation":[{"name":"Department of Computer and Information Security Management, Fujian Police College, Fuzhou, Fujian, China"},{"name":"Collaborative Innovation Research Center of Intelligent Policing, Fujian Police College, Fuzhou, Fujian, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenxi","family":"Liu","sequence":"additional","affiliation":[{"name":"Collaborative Innovation Research Center of Intelligent Policing, Fujian Police College, Fuzhou, Fujian, China"},{"name":"College of Computer and Data Science, Fuzhou University, Fuzhou, Fujian, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"4443","published-online":{"date-parts":[[2026,3,9]]},"reference":[{"key":"10.7717\/peerj-cs.3667\/ref-1","article-title":"Semantic annotation and transcoding of soccer videos","author":"Bertini","year":"2004"},{"key":"10.7717\/peerj-cs.3667\/ref-2","doi-asserted-by":"crossref","DOI":"10.1109\/ICCVW.2019.00246","article-title":"GCNet: non-local networks meet squeeze-excitation networks and beyond","author":"Cao","year":"2019"},{"key":"10.7717\/peerj-cs.3667\/ref-3","first-page":"213","article-title":"End-to-end object detection with transformers","author":"Carion","year":"2020"},{"key":"10.7717\/peerj-cs.3667\/ref-4","first-page":"93","article-title":"A graph-based method for soccer action spotting using unsupervised player classification","author":"Cartas","year":"2022"},{"key":"10.7717\/peerj-cs.3667\/ref-5","first-page":"10337","article-title":"Memory enhanced global-local aggregation for video object detection","author":"Chen","year":"2020"},{"key":"10.7717\/peerj-cs.3667\/ref-6","first-page":"8126","article-title":"Transformer tracking","author":"Chen","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-7","first-page":"13126","article-title":"A context-aware loss function for action spotting in soccer videos","author":"Cioppa","year":"2020"},{"key":"10.7717\/peerj-cs.3667\/ref-8","first-page":"4508","article-title":"SoccerNet-v2: a dataset and benchmarks for holistic understanding of broadcast soccer videos","author":"Deliege","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-9","article-title":"An image is worth 16x16 words: transformers for image recognition at scale","author":"Dosovitskiy","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-10","first-page":"3146","article-title":"Dual attention network for scene segmentation","author":"Fu","year":"2019"},{"key":"10.7717\/peerj-cs.3667\/ref-11","first-page":"1711","article-title":"SoccerNet: a scalable dataset for action spotting in soccer videos","author":"Giancola","year":"2018"},{"key":"10.7717\/peerj-cs.3667\/ref-12","first-page":"4490","article-title":"Temporally-aware feature pooling for action spotting in soccer broadcasts","author":"Giancola","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-13","first-page":"244","article-title":"Video action transformer network","author":"Girdhar","year":"2019"},{"key":"10.7717\/peerj-cs.3667\/ref-14","first-page":"7231","article-title":"Mining contextual information beyond image for semantic segmentation","author":"Jin","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-15","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1705.06950","article-title":"The kinetics human action video dataset","author":"Kay","year":"2017"},{"key":"10.7717\/peerj-cs.3667\/ref-16","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2102.04432","article-title":"Colorization transformer","author":"Kumar","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-17","first-page":"3054","article-title":"Video prediction recalling long-term motion context via memory alignment learning","author":"Lee","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-18","first-page":"6102","article-title":"Memory-based neighbourhood embedding for visual recognition","author":"Li","year":"2019"},{"key":"10.7717\/peerj-cs.3667\/ref-19","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2103.15320","article-title":"TFPose: direct human pose estimation with transformers","author":"Mao","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-20","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2101.08540","article-title":"Activity graph transformer for temporal action localization","author":"Nawhal","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-21","first-page":"9226","article-title":"Video object segmentation using space-time memory networks","author":"Oh","year":"2019"},{"issue":"1","key":"10.7717\/peerj-cs.3667\/ref-22","doi-asserted-by":"publisher","first-page":"105","DOI":"10.1587\/transinf.2022edp7210","article-title":"Efficient action spotting using saliency feature weighting","volume":"107","author":"Shi","year":"2024","journal-title":"IEICE Transactions on Information and Systems"},{"key":"10.7717\/peerj-cs.3667\/ref-23","first-page":"13526","article-title":"Relaxed transformer decoders for direct action proposal generation","author":"Tan","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-24","first-page":"7699","article-title":"RMS-NET: regression and masking for soccer event spotting","author":"Tomei","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-25","first-page":"5552","article-title":"Video classification with channel-separated convolutional networks","author":"Tran","year":"2019"},{"key":"10.7717\/peerj-cs.3667\/ref-26","article-title":"Attention is all you need","author":"Vaswani","year":"2017"},{"key":"10.7717\/peerj-cs.3667\/ref-27","first-page":"7794","article-title":"Non-local neural networks","author":"Wang","year":"2018"},{"key":"10.7717\/peerj-cs.3667\/ref-28","first-page":"284","article-title":"Long-term feature banks for detailed video understanding","author":"Wu","year":"2019"},{"key":"10.7717\/peerj-cs.3667\/ref-29","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2103.15145","article-title":"Transcenter: transformers with dense queries for multiple-object tracking","author":"Xu","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-30","first-page":"591","article-title":"Temporal pyramid network for action recognition","author":"Yang","year":"2020"},{"key":"10.7717\/peerj-cs.3667\/ref-31","doi-asserted-by":"publisher","first-page":"173","DOI":"10.1007\/978-3-030-58539-6_11","article-title":"Object-contextual representations for semantic segmentation","volume":"12351","author":"Yuan","year":"2020"},{"key":"10.7717\/peerj-cs.3667\/ref-32","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1809.00916","article-title":"OCNet: object context network for scene parsing","author":"Yuan","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-33","first-page":"267","article-title":"PSANet: point-wise spatial attention network for scene parsing","author":"Zhao","year":"2018"},{"key":"10.7717\/peerj-cs.3667\/ref-34","first-page":"598","article-title":"Invariance matters: exemplar memory for domain adaptive person re-identification","author":"Zhong","year":"2019"},{"key":"10.7717\/peerj-cs.3667\/ref-35","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2106.14447","article-title":"Feature combination meets attention: baidu soccer embeddings and transformer based temporal detection","author":"Zhou","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-36","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2103.14829","article-title":"Looking beyond two frames: end-to-end multi-object tracking using spatial and temporal transformers","author":"Zhu","year":"2021"},{"key":"10.7717\/peerj-cs.3667\/ref-37","first-page":"103","article-title":"A transformer-based system for action spotting in soccer videos","author":"Zhu","year":"2022"},{"key":"10.7717\/peerj-cs.3667\/ref-38","first-page":"593","article-title":"Asymmetric non-local neural networks for semantic segmentation","author":"Zhu","year":"2019"}],"container-title":["PeerJ Computer Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/peerj.com\/articles\/cs-3667.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/peerj.com\/articles\/cs-3667.xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/peerj.com\/articles\/cs-3667.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/peerj.com\/articles\/cs-3667.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,9]],"date-time":"2026-03-09T08:20:31Z","timestamp":1773044431000},"score":1,"resource":{"primary":{"URL":"https:\/\/peerj.com\/articles\/cs-3667"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,9]]},"references-count":38,"alternative-id":["10.7717\/peerj-cs.3667"],"URL":"https:\/\/doi.org\/10.7717\/peerj-cs.3667","archive":["CLOCKSS","LOCKSS","Portico"],"relation":{},"ISSN":["2376-5992"],"issn-type":[{"value":"2376-5992","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,9]]},"article-number":"e3667"}}