{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T15:01:47Z","timestamp":1784905307681,"version":"3.55.0"},"reference-count":49,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2023,7,27]],"date-time":"2023-07-27T00:00:00Z","timestamp":1690416000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,7,27]],"date-time":"2023-07-27T00:00:00Z","timestamp":1690416000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimedia Systems"],"published-print":{"date-parts":[[2023,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Point-level temporal action localization (PTAL) aims to locate action instances in untrimmed videos with only one timestamp annotation for each action instance. Existing methods adopt the localization-by-classification paradigm to locate action boundaries in the temporal class activation map (TCAM) by thresholding, also known as TCAM-based method. However, TCAM-based methods are limited by the gap between classification and localization tasks, since TCAM is generated by a classification network. To address this issue, we propose a re-training framework for the PTAL task, also known as LPR. This framework consists of two stages: pseudo-label generation and re-training. In the pseudo-label generation stage, we propose a feature embedding module based on a transformer encoder to capture global context features and optimize pseudo-labels\u2019 quality by leveraging point-level annotations. In the re-training stage, LPR uses the above pseudo-labels as supervision to locate action instances with a temporal action localization network rather than generating TCAMs. Furthermore, to alleviate the effects of label noise in the pseudo-labels, we propose a joint learning classification module (JLCM) in the re-training stage. This module contains two classification sub-modules that simultaneously predict action categories and are guided by a jointly determined clean set for network training. The proposed framework achieves state-of-the-art localization performance on both the THUMOS\u201914 and BEOID datasets.<\/jats:p>","DOI":"10.1007\/s00530-023-01128-4","type":"journal-article","created":{"date-parts":[[2023,7,27]],"date-time":"2023-07-27T12:02:23Z","timestamp":1690459343000},"page":"2545-2562","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["LPR: learning point-level temporal action localization through re-training"],"prefix":"10.1007","volume":"29","author":[{"given":"Zhenying","family":"Fang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianping","family":"Fan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun","family":"Yu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,7,27]]},"reference":[{"issue":"11","key":"1128_CR1","doi-asserted-by":"publisher","first-page":"1838","DOI":"10.1109\/JPROC.2021.3117472","volume":"109","author":"E Apostolidis","year":"2021","unstructured":"Apostolidis, E., Adamantidou, E., Metsai, A.I., Mezaris, V., Patras, I.: Video summarization using deep neural networks: a survey. Proc. IEEE 109(11), 1838\u20131863 (2021)","journal-title":"Proc. IEEE"},{"key":"1128_CR2","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-023-01071-4","author":"S Derdiyok","year":"2023","unstructured":"Derdiyok, S., Patlar Akbulut, F.: Biosignal based emotion-oriented video summarization. Multimed. Syst. (2023). https:\/\/doi.org\/10.1007\/s00530-023-01071-4","journal-title":"Multimed. Syst."},{"key":"1128_CR3","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-015-0457-6","author":"M-C Yeh","year":"2015","unstructured":"Yeh, M.-C., Tsai, Y.-W., Hsu, H.-C.: A content-based approach for detecting highlights in action movies. Multimed. Syst. (2015). https:\/\/doi.org\/10.1007\/s00530-015-0457-6","journal-title":"Multimed. Syst."},{"key":"1128_CR4","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-022-00978-8","author":"A Khan","year":"2022","unstructured":"Khan, A., Rao, Y., Shao, J.: Enet: event based highlight generation network for broadcast sports videos. Multimed. Syst. (2022). https:\/\/doi.org\/10.1007\/s00530-022-00978-8","journal-title":"Multimed. Syst."},{"key":"1128_CR5","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-020-00684-3","author":"H Shingrakhia","year":"2020","unstructured":"Shingrakhia, H., Patel, H.: Emperor penguin optimized event recognition and summarization for cricket highlight generation. Multimed. Syst. (2020). https:\/\/doi.org\/10.1007\/s00530-020-00684-3","journal-title":"Multimed. Syst."},{"key":"1128_CR6","doi-asserted-by":"crossref","unstructured":"Sultani, W., Chen, C., Shah, M.: Real-world anomaly detection in surveillance videos. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6479\u20136488 (2018)","DOI":"10.1109\/CVPR.2018.00678"},{"key":"1128_CR7","doi-asserted-by":"crossref","unstructured":"Liu, M., Wang, X., Nie, L., Tian, Q., Chen, B., Chua, T.-S.: Cross-modal moment localization in videos. In: Proceedings of the 26th ACM International Conference on Multimedia, pp. 843\u2013851 (2018)","DOI":"10.1145\/3240508.3240549"},{"issue":"9","key":"1128_CR8","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3560815","volume":"55","author":"M Liu","year":"2023","unstructured":"Liu, M., Nie, L., Wang, Y., Wang, M., Rui, Y.: A survey on video moment localization. ACM Comput. Surv. 55(9), 1\u201337 (2023)","journal-title":"ACM Comput. Surv."},{"key":"1128_CR9","doi-asserted-by":"crossref","unstructured":"Zhao, Y., Xiong, Y., Wang, L., Wu, Z., Tang, X., Lin, D.: Temporal action detection with structured segment networks. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 2914\u20132923 (2017)","DOI":"10.1109\/ICCV.2017.317"},{"key":"1128_CR10","doi-asserted-by":"crossref","unstructured":"Lin, T., Zhao, X., Su, H., Wang, C., Yang, M.: Bsn: boundary sensitive network for temporal action proposal generation. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 3\u201319 (2018)","DOI":"10.1007\/978-3-030-01225-0_1"},{"key":"1128_CR11","doi-asserted-by":"crossref","unstructured":"Xu, M., Zhao, C., Rojas, D.S., Thabet, A., Ghanem, B.: G-tad: sub-graph localization for temporal action detection. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 10156\u201310165 (2020)","DOI":"10.1109\/CVPR42600.2020.01017"},{"key":"1128_CR12","doi-asserted-by":"crossref","unstructured":"Zhang, C.-L., Wu, J., Li, Y.: Actionformer: localizing moments of actions with transformers. In: European Conference on Computer Vision, pp. 492\u2013510 (2022). Springer","DOI":"10.1007\/978-3-031-19772-7_29"},{"key":"1128_CR13","doi-asserted-by":"crossref","unstructured":"Ma, F., Zhu, L., Yang, Y., Zha, S., Kundu, G., Feiszli, M., Shou, Z.: Sf-net: single-frame supervision for temporal action localization. In: Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part IV 16, pp. 420\u2013437 (2020). Springer","DOI":"10.1007\/978-3-030-58548-8_25"},{"key":"1128_CR14","doi-asserted-by":"crossref","unstructured":"Lee, P., Byun, H.: Learning action completeness from points for weakly-supervised temporal action localization. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 13648\u201313657 (2021)","DOI":"10.1109\/ICCV48922.2021.01339"},{"key":"1128_CR15","doi-asserted-by":"publisher","first-page":"7363","DOI":"10.1109\/TIP.2022.3222623","volume":"31","author":"J Fu","year":"2022","unstructured":"Fu, J., Gao, J., Xu, C.: Compact representation and reliable classification learning for point-level weakly-supervised action localization. IEEE Trans. Image Process. 31, 7363\u20137377 (2022)","journal-title":"IEEE Trans. Image Process."},{"key":"1128_CR16","doi-asserted-by":"crossref","unstructured":"Ju, C., Zhao, P., Chen, S., Zhang, Y., Wang, Y., Tian, Q.: Divide and conquer for single-frame temporal action localization. In: Proceedings of the IEEE\/CVF International Conference on Computer Vision, pp. 13455\u201313464 (2021)","DOI":"10.1109\/ICCV48922.2021.01320"},{"issue":"12","key":"1128_CR17","doi-asserted-by":"publisher","first-page":"9814","DOI":"10.1109\/TPAMI.2021.3132058","volume":"44","author":"L Yang","year":"2021","unstructured":"Yang, L., Han, J., Zhao, T., Lin, T., Zhang, D., Chen, J.: Background-click supervision for temporal action localization. IEEE Trans. Pattern Anal. Mach. Intell. 44(12), 9814\u20139829 (2021)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"3","key":"1128_CR18","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1145\/3446776","volume":"64","author":"C Zhang","year":"2021","unstructured":"Zhang, C., Bengio, S., Hardt, M., Recht, B., Vinyals, O.: Understanding deep learning (still) requires rethinking generalization. Commun. ACM 64(3), 107\u2013115 (2021)","journal-title":"Commun. ACM"},{"key":"1128_CR19","doi-asserted-by":"crossref","unstructured":"Xu, H., Das, A., Saenko, K.: R-c3d: region convolutional 3d network for temporal activity detection. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 5783\u20135792 (2017)","DOI":"10.1109\/ICCV.2017.617"},{"key":"1128_CR20","doi-asserted-by":"crossref","unstructured":"Lin, C., Xu, C., Luo, D., Wang, Y., Tai, Y., Wang, C., Li, J., Huang, F., Fu, Y.: Learning salient boundary feature for anchor-free temporal action localization. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 3320\u20133329 (2021)","DOI":"10.1109\/CVPR46437.2021.00333"},{"key":"1128_CR21","doi-asserted-by":"crossref","unstructured":"Fang, Z., Zhu, S., Yu, J., Tian, Q.: Pcpcad: proposal complementary action detector. In: 2019 IEEE International Conference on Multimedia and Expo (ICME), pp. 424\u2013429 (2019). IEEE","DOI":"10.1109\/ICME.2019.00080"},{"key":"1128_CR22","doi-asserted-by":"crossref","unstructured":"Liu, Y., Ma, L., Zhang, Y., Liu, W., Chang, S.-F.: Multi-granularity generator for temporal action proposal. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 3604\u20133613 (2019)","DOI":"10.1109\/CVPR.2019.00372"},{"key":"1128_CR23","doi-asserted-by":"crossref","unstructured":"Cheng, F., Bertasius, G.: Tallformer: temporal action localization with a long-memory transformer. In: Computer Vision\u2013ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23\u201327, 2022, Proceedings, Part XXXIV, pp. 503\u2013521 (2022). Springer","DOI":"10.1007\/978-3-031-19830-4_29"},{"key":"1128_CR24","doi-asserted-by":"crossref","unstructured":"Wang, L., Xiong, Y., Lin, D., Van\u00a0Gool, L.: Untrimmednets for weakly supervised action recognition and detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4325\u20134334 (2017)","DOI":"10.1109\/CVPR.2017.678"},{"key":"1128_CR25","doi-asserted-by":"crossref","unstructured":"Nguyen, P., Liu, T., Prasad, G., Han, B.: Weakly supervised action localization by sparse temporal pooling network. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6752\u20136761 (2018)","DOI":"10.1109\/CVPR.2018.00706"},{"key":"1128_CR26","doi-asserted-by":"crossref","unstructured":"Liu, D., Jiang, T., Wang, Y.: Completeness modeling and context separation for weakly supervised temporal action localization. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 1298\u20131307 (2019)","DOI":"10.1109\/CVPR.2019.00139"},{"key":"1128_CR27","doi-asserted-by":"crossref","unstructured":"Lee, P., Uh, Y., Byun, H.: Background suppression network for weakly-supervised temporal action localization. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, pp. 11320\u201311327 (2020)","DOI":"10.1609\/aaai.v34i07.6793"},{"key":"1128_CR28","doi-asserted-by":"crossref","unstructured":"He, B., Yang, X., Kang, L., Cheng, Z., Zhou, X., Shrivastava, A.: Asm-loc: action-aware segment modeling for weakly-supervised temporal action localization. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 13925\u201313935 (2022)","DOI":"10.1109\/CVPR52688.2022.01355"},{"issue":"4","key":"1128_CR29","doi-asserted-by":"publisher","first-page":"1529","DOI":"10.1007\/s00530-022-00912-y","volume":"28","author":"H Xia","year":"2022","unstructured":"Xia, H., Zhan, Y., Cheng, K.: Spatial-temporal correlations learning and action-background jointed attention for weakly-supervised temporal action localization. Multimed. Syst. 28(4), 1529\u20131541 (2022)","journal-title":"Multimed. Syst."},{"key":"1128_CR30","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2022.118965","volume":"213","author":"P Li","year":"2023","unstructured":"Li, P., Cao, J., Ye, X.: Prototype contrastive learning for point-supervised temporal action detection. Expert Syst. Appl. 213, 118965 (2023)","journal-title":"Expert Syst. Appl."},{"key":"1128_CR31","doi-asserted-by":"crossref","unstructured":"Qin, J., Wu, J., Xiao, X., Li, L., Wang, X.: Activation modulation and recalibration scheme for weakly supervised semantic segmentation. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, pp. 2117\u20132125 (2022)","DOI":"10.1609\/aaai.v36i2.20108"},{"key":"1128_CR32","doi-asserted-by":"crossref","unstructured":"Lee, J., Kim, E., Yoon, S.: Anti-adversarially manipulated attributions for weakly and semi-supervised semantic segmentation. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 4071\u20134080 (2021)","DOI":"10.1109\/CVPR46437.2021.00406"},{"key":"1128_CR33","unstructured":"Ju, C., Zhao, P., Zhang, Y., Wang, Y., Tian, Q.: Point-level temporal action localization: bridging fully-supervised proposals to weakly-supervised losses (2020). arXiv:2012.08236"},{"key":"1128_CR34","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., Polosukhin, I.: Attention is all you need. Advances in Neural Information Processing Systems, vol. 30 (2017)"},{"key":"1128_CR35","unstructured":"Ba, J.L., Kiros, J.R., Hinton, G.E.: Layer normalization (2016). arXiv:1607.06450"},{"key":"1128_CR36","first-page":"30392","volume":"34","author":"T Xiao","year":"2021","unstructured":"Xiao, T., Singh, M., Mintun, E., Darrell, T., Doll\u00e1r, P., Girshick, R.: Early convolutions help transformers see better. Adv. Neural Inf. Process. Syst. 34, 30392\u201330400 (2021)","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"1128_CR37","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Goyal, P., Girshick, R., He, K., Doll\u00e1r, P.: Focal loss for dense object detection. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 2980\u20132988 (2017)","DOI":"10.1109\/ICCV.2017.324"},{"key":"1128_CR38","doi-asserted-by":"crossref","unstructured":"Wei, H., Feng, L., Chen, X., An, B.: Combating noisy labels by agreement: a joint training method with co-regularization. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 13726\u201313735 (2020)","DOI":"10.1109\/CVPR42600.2020.01374"},{"key":"1128_CR39","doi-asserted-by":"crossref","unstructured":"Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., Ren, D.: Distance-iou loss: faster and better learning for bounding box regression. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, pp. 12993\u201313000 (2020)","DOI":"10.1609\/aaai.v34i07.6999"},{"key":"1128_CR40","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.cviu.2016.10.018","volume":"155","author":"H Idrees","year":"2017","unstructured":"Idrees, H., Zamir, A.R., Jiang, Y.-G., Gorban, A., Laptev, I., Sukthankar, R., Shah, M.: The thumos challenge on action recognition for videos \u201cin the wild\u2019\u2019. Comput. Vis. Image Underst. 155, 1\u201323 (2017)","journal-title":"Comput. Vis. Image Underst."},{"key":"1128_CR41","doi-asserted-by":"crossref","unstructured":"Calway, A., Mayol-Cuevas, W., Damen, D., Haines, O., Leelasawassuk, T.: Discovering task relevant objects and their modes of interaction from multi-user egocentric video. In: BMVC, pp. 30\u20131 (2015)","DOI":"10.5244\/C.28.30"},{"key":"1128_CR42","doi-asserted-by":"crossref","unstructured":"Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6299\u20136308 (2017)","DOI":"10.1109\/CVPR.2017.502"},{"key":"1128_CR43","doi-asserted-by":"crossref","unstructured":"Liu, X., Bai, S., Bai, X.: An empirical study of end-to-end temporal action detection. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 20010\u201320019 (2022)","DOI":"10.1109\/CVPR52688.2022.01938"},{"key":"1128_CR44","doi-asserted-by":"crossref","unstructured":"Li, J., Yang, T., Ji, W., Wang, J., Cheng, L.: Exploring denoised cross-video contrast for weakly-supervised temporal action localization. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 19914\u201319924 (2022)","DOI":"10.1109\/CVPR52688.2022.01929"},{"key":"1128_CR45","doi-asserted-by":"crossref","unstructured":"Zhao, Y., Zhang, H., Gao, Z., Gao, W., Wang, M., Chen, S.: A novel action saliency and context-aware network for weakly-supervised temporal action localization. IEEE Trans. Multimed. (2023)","DOI":"10.1109\/TMM.2023.3234362"},{"key":"1128_CR46","doi-asserted-by":"crossref","unstructured":"Zhou, J., Wu, Y.: Temporal feature enhancement dilated convolution network for weakly-supervised temporal action localization. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, pp. 6028\u20136037 (2023)","DOI":"10.1109\/WACV56688.2023.00597"},{"key":"1128_CR47","doi-asserted-by":"crossref","unstructured":"Moltisanti, D., Fidler, S., Damen, D.: Action recognition from single timestamp supervision in untrimmed videos. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 9915\u20139924 (2019)","DOI":"10.1109\/CVPR.2019.01015"},{"key":"1128_CR48","unstructured":"Ju, C., Zhao, P., Zhang, Y., Wang, Y., Tian, Q.: Point-level temporal action localization: bridging fully-supervised proposals to weakly-supervised losses (2020). arXiv:2012.08236"},{"key":"1128_CR49","doi-asserted-by":"crossref","unstructured":"Alwassel, H., Heilbron, F.C., Escorcia, V., Ghanem, B.: Diagnosing error in temporal action detectors. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 256\u2013272 (2018)","DOI":"10.1007\/978-3-030-01219-9_16"}],"container-title":["Multimedia Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00530-023-01128-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00530-023-01128-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00530-023-01128-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,25]],"date-time":"2024-10-25T04:27:23Z","timestamp":1729830443000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00530-023-01128-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,27]]},"references-count":49,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2023,10]]}},"alternative-id":["1128"],"URL":"https:\/\/doi.org\/10.1007\/s00530-023-01128-4","relation":{},"ISSN":["0942-4962","1432-1882"],"issn-type":[{"value":"0942-4962","type":"print"},{"value":"1432-1882","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,27]]},"assertion":[{"value":"27 May 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 June 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 July 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}