{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,25]],"date-time":"2026-07-25T16:21:59Z","timestamp":1784996519789,"version":"3.55.0"},"reference-count":128,"publisher":"Springer Science and Business Media LLC","issue":"12","license":[{"start":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T00:00:00Z","timestamp":1760659200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T00:00:00Z","timestamp":1760659200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Human action recognition (HAR) encompasses the task of monitoring human activities across various domains, including but not limited to medical, educational, entertainment, visual surveillance, video retrieval, and the identification of anomalous activities. Over the past decade, the field of HAR has witnessed substantial progress by leveraging convolutional neural networks (CNNs) and recurrent neural networks (RNNs) to effectively extract and comprehend intricate information, thereby enhancing the overall performance of HAR systems. Recently, the domain of computer vision has witnessed the emergence of Vision Transformers (ViTs) as a potent solution. The efficacy of Transformer architecture has been validated beyond the confines of image analysis, extending their applicability to diverse video-related tasks. Notably, within this landscape, the research community has shown keen interest in HAR, acknowledging its manifold utility and widespread adoption across various domains. However, HAR remains a challenging task due to variations in human motion, occlusions, viewpoint differences, background clutter, and the need for efficient spatio-temporal feature extraction. Additionally, the trade-off between computational efficiency and recognition accuracy remains a significant obstacle, particularly with the adoption of deep learning models requiring extensive training data and resources. This article aims to present an encompassing survey that focuses on CNNs and the evolution of RNNs to ViTs given their importance in the domain of HAR. By conducting a thorough examination of existing literature and exploring emerging trends, this study undertakes a critical analysis and synthesis of the accumulated knowledge in this field. Additionally, it investigates the ongoing efforts to develop hybrid approaches. Following this direction, this article presents a novel hybrid model that seeks to integrate the inherent strengths of CNNs and ViTs.<\/jats:p>","DOI":"10.1007\/s10462-025-11388-3","type":"journal-article","created":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T02:55:48Z","timestamp":1760669748000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":48,"title":["CNNs, RNNs and Transformers in human action recognition: a survey and a hybrid model"],"prefix":"10.1007","volume":"58","author":[{"given":"Khaled","family":"Alomar","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Halil Ibrahim","family":"Aysel","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaohao","family":"Cai","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,10,17]]},"reference":[{"key":"11388_CR1","doi-asserted-by":"crossref","unstructured":"Abdelbaky A, Aly S (2020) Human action recognition based on simple deep convolution network Pcanet. In 2020 International conference on innovative trends in communication and computer engineering (ITCE). IEEE, pp. 257\u2013262","DOI":"10.1109\/ITCE48509.2020.9047769"},{"key":"11388_CR2","doi-asserted-by":"crossref","unstructured":"Ahmadabadi H, Manzari ON, Ayatollahi A (2023) Distilling knowledge from CNN-transformer models for enhanced human action recognition. In 2023 13th International conference on computer and knowledge engineering (ICCKE). IEEE, pp. 180\u2013184","DOI":"10.1109\/ICCKE60553.2023.10326272"},{"key":"11388_CR3","doi-asserted-by":"crossref","unstructured":"Alomar K, Cai X (2023) TransNet: A transfer learning-based network for human action recognition. In 2023 International conference on machine learning and applications (ICMLA). IEEE, pp. 1825\u20131832","DOI":"10.1109\/ICMLA58977.2023.00277"},{"key":"11388_CR4","doi-asserted-by":"crossref","unstructured":"Arnab A, Dehghani M, Heigold G, Sun C, Lu\u010di\u0107 M, Schmid C (2021) Vivit: a video vision transformer. In Proceedings of the IEEE\/CVF international conference on computer vision, pp. 6836\u20136846","DOI":"10.1109\/ICCV48922.2021.00676"},{"key":"11388_CR5","doi-asserted-by":"publisher","first-page":"471","DOI":"10.1016\/j.procs.2018.07.059","volume":"133","author":"J Arunnehru","year":"2018","unstructured":"Arunnehru J, Chamundeeswari G, Bharathi SP (2018) Human action recognition using 3d convolutional neural networks with 3d motion cuboids in surveillance videos. Procedia Comput Sci 133:471\u2013477","journal-title":"Procedia Comput Sci"},{"key":"11388_CR6","unstructured":"Bahdanau D, Cho K, Bengio Y (2014) Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473"},{"key":"11388_CR7","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1007\/BF01420984","volume":"12","author":"JL Barron","year":"1994","unstructured":"Barron JL, Fleet DJ, Beauchemin SS (1994) Performance of optical flow techniques. Int J Comput Vis 12:43\u201377","journal-title":"Int J Comput Vis"},{"issue":"28","key":"11388_CR8","doi-asserted-by":"publisher","first-page":"40431","DOI":"10.1007\/s11042-022-12856-6","volume":"81","author":"SS Basha","year":"2022","unstructured":"Basha SS, Pulabaigari V, Mukherjee S (2022) An information-rich sampling technique over spatio-temporal cnn for classification of human actions in videos. Multimedia Tools Appl 81(28):40431\u201340449","journal-title":"Multimedia Tools Appl"},{"issue":"2","key":"11388_CR9","doi-asserted-by":"publisher","first-page":"157","DOI":"10.1109\/72.279181","volume":"5","author":"Y Bengio","year":"1994","unstructured":"Bengio Y, Simard P, Frasconi P (1994) Learning long-term dependencies with gradient descent is difficult. IEEE Trans Neural Netw 5(2):157\u2013166","journal-title":"IEEE Trans Neural Netw"},{"key":"11388_CR10","first-page":"4","volume":"2","author":"G Bertasius","year":"2021","unstructured":"Bertasius G, Wang H, Torresani L (2021) Is space-time attention all you need for video understanding? ICML 2:4","journal-title":"ICML"},{"issue":"4","key":"11388_CR11","doi-asserted-by":"publisher","first-page":"3279","DOI":"10.1109\/TKDE.2021.3126456","volume":"35","author":"G Brauwers","year":"2021","unstructured":"Brauwers G, Frasincar F (2021) A general survey on attention mechanisms in deep learning. IEEE Trans Knowl Data Eng 35(4):3279\u20133298","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"11388_CR12","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A et al (2020) Language models are few-shot learners. Adv Neural Inf Process Syst 33:1877\u20131901","journal-title":"Adv Neural Inf Process Syst"},{"key":"11388_CR13","doi-asserted-by":"crossref","unstructured":"Caba\u00a0Heilbron F, Escorcia V, Ghanem B, Carlos\u00a0Niebles J (2015) Activitynet: a large-scale video benchmark for human activity understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 961\u2013970","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"11388_CR14","doi-asserted-by":"crossref","unstructured":"Carion N, Massa F, Synnaeve G, Usunier N, Kirillov A, Zagoruyko S (2020) End-to-end object detection with transformers. In European conference on computer vision. Springer, pp. 213\u2013229","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"11388_CR15","doi-asserted-by":"crossref","unstructured":"Carreira J, Zisserman A (2017) Quo Vadis, action recognition? A new model and the kinetics dataset. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6299\u20136308","DOI":"10.1109\/CVPR.2017.502"},{"key":"11388_CR17","doi-asserted-by":"crossref","unstructured":"Chen J, Ho CM (2022) MM-ViT: Multi-modal video transformer for compressed video action recognition. In Proceedings of the IEEE\/CVF winter conference on applications of computer vision, pp. 1910\u20131921","DOI":"10.1109\/WACV51458.2022.00086"},{"issue":"8","key":"11388_CR16","doi-asserted-by":"publisher","first-page":"11109","DOI":"10.1007\/s11063-023-11367-1","volume":"55","author":"T Chen","year":"2023","unstructured":"Chen T, Mo L (2023) Swin-fusion: swin-transformer with feature fusion for human action recognition. Neural Process Lett 55(8):11109\u201311130","journal-title":"Neural Process Lett"},{"key":"11388_CR18","doi-asserted-by":"crossref","unstructured":"Cho K, Van\u00a0Merri\u00ebnboer B, Gulcehre C, Bahdanau D, Bougares F, Schwenk H, Bengio Y (2014) Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078","DOI":"10.3115\/v1\/D14-1179"},{"key":"11388_CR19","unstructured":"Chung J, Gulcehre C, Cho K, Bengio Y (2014) Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555"},{"key":"11388_CR20","doi-asserted-by":"publisher","first-page":"273","DOI":"10.1023\/A:1022627411411","volume":"20","author":"C Cortes","year":"1995","unstructured":"Cortes C, Vapnik V (1995) Support-vector networks. Mach Learn 20:273\u2013297","journal-title":"Mach Learn"},{"key":"11388_CR21","doi-asserted-by":"crossref","unstructured":"Cosmin\u00a0Duta I, Ionescu B, Aizawa K, Sebe N (2017) Spatio-temporal vector of locally max pooled features for action recognition in videos. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3097\u20133106","DOI":"10.1109\/CVPR.2017.341"},{"key":"11388_CR22","doi-asserted-by":"crossref","unstructured":"Dalal N, Triggs B (2005) Histograms of oriented gradients for human detection. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR\u201905). IEEE, Volume\u00a01, pp. 886\u2013893","DOI":"10.1109\/CVPR.2005.177"},{"key":"11388_CR23","doi-asserted-by":"crossref","unstructured":"Dar G, Geva M, Gupta A, Berant J (2022) Analyzing transformers in embedding space. arXiv preprint arXiv:2209.02535","DOI":"10.18653\/v1\/2023.acl-long.893"},{"key":"11388_CR24","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li LJ, Li K, Fei-Fei L (2009) Imagenet: a large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition. IEEE, pp. 248\u2013255","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"11388_CR25","doi-asserted-by":"crossref","unstructured":"Devlin J, Chang MW, Lee K, Toutanova K (2019) BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, Volume 1 (Long and Short Papers), pp. 4171\u20134186","DOI":"10.18653\/v1\/N19-1423"},{"key":"11388_CR26","volume-title":"Probability and statistics","author":"JL Devore","year":"2000","unstructured":"Devore JL (2000) Probability and statistics. Brooks\/Cole, Pacific Grove"},{"key":"11388_CR27","unstructured":"Diba A, Fayyaz M, Sharma V, Karami AH, Arzani MM, Yousefzadeh R, Van\u00a0Gool L (2017) Temporal 3d convnets: new architecture and transfer learning for video classification. arXiv preprint arXiv:1711.08200"},{"key":"11388_CR28","doi-asserted-by":"crossref","unstructured":"Djenouri Y, Belbachir AN (2023) A hybrid visual transformer for efficient deep human activity recognition. In Proceedings of the IEEE\/CVF international conference on computer vision, pp. 721\u2013730","DOI":"10.1109\/ICCVW60793.2023.00080"},{"key":"11388_CR29","doi-asserted-by":"crossref","unstructured":"Donahue J, Anne\u00a0Hendricks L, Guadarrama S, Rohrbach M,\u00a0Venugopalan S,\u00a0Saenko K,\u00a0Darrell T (2015) Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2625\u20132634","DOI":"10.1109\/CVPR.2015.7298878"},{"key":"11388_CR30","unstructured":"Dosovitskiy A, Beyer L,\u00a0Kolesnikov A,\u00a0Weissenborn D, Zhai X,\u00a0Unterthiner T,\u00a0Dehghani M,\u00a0Minderer M,\u00a0Heigold G,\u00a0Gelly S et\u00a0al (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929"},{"key":"11388_CR31","doi-asserted-by":"crossref","unstructured":"Fan H, Xiong B, Mangalam K, Li Y, Yan Z, Malik J, Feichtenhofer C (2021) Multiscale vision transformers. In Proceedings of the IEEE\/CVF international conference on computer vision, pp. 6824\u20136835","DOI":"10.1109\/ICCV48922.2021.00675"},{"key":"11388_CR32","doi-asserted-by":"crossref","unstructured":"Feichtenhofer C (2020) X3d: Expanding architectures for efficient video recognition. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp. 203\u2013213","DOI":"10.1109\/CVPR42600.2020.00028"},{"key":"11388_CR34","doi-asserted-by":"crossref","unstructured":"Feichtenhofer C, Pinz A, Zisserman A (2016) Convolutional two-stream network fusion for video action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1933\u20131941","DOI":"10.1109\/CVPR.2016.213"},{"key":"11388_CR33","doi-asserted-by":"crossref","unstructured":"Feichtenhofer C, Fan H, Malik J, He K (2019) Slowfast networks for video recognition. In Proceedings of the IEEE\/CVF international conference on computer vision, pp. 6202\u20136211","DOI":"10.1109\/ICCV.2019.00630"},{"issue":"4","key":"11388_CR35","doi-asserted-by":"publisher","first-page":"193","DOI":"10.1007\/BF00344251","volume":"36","author":"K Fukushima","year":"1980","unstructured":"Fukushima K (1980) Neocognitron: a self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position. Biol Cybern 36(4):193\u2013202","journal-title":"Biol Cybern"},{"key":"11388_CR36","doi-asserted-by":"crossref","unstructured":"Geng C, Song J (2016) Human action recognition based on convolutional neural networks with a convolutional auto-encoder. In 2015 5th international conference on computer sciences and automation engineering (ICCSAE 2015). Atlantis Press, pp. 933\u2013938","DOI":"10.2991\/iccsae-15.2016.173"},{"issue":"10","key":"11388_CR37","doi-asserted-by":"publisher","first-page":"2451","DOI":"10.1162\/089976600300015015","volume":"12","author":"FA Gers","year":"2000","unstructured":"Gers FA, Schmidhuber J, Cummins F (2000) Learning to forget: continual prediction with lstm. Neural Comput 12(10):2451\u20132471","journal-title":"Neural Comput"},{"key":"11388_CR38","doi-asserted-by":"crossref","unstructured":"Ghadiyaram D, Tran D, Mahajan D (2019) Large-scale weakly-supervised pre-training for video action recognition. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp. 12046\u201312055","DOI":"10.1109\/CVPR.2019.01232"},{"key":"11388_CR39","doi-asserted-by":"crossref","unstructured":"Gong G, Wang X, Mu Y, Tian Q (2020) Learning temporal co-attention models for unsupervised video action localization. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp. 9819\u20139828","DOI":"10.1109\/CVPR42600.2020.00984"},{"key":"11388_CR40","unstructured":"Graves A (2013) Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850"},{"key":"11388_CR41","doi-asserted-by":"crossref","unstructured":"Graves A, Mohamed AR, Hinton G (2013) Speech recognition with deep recurrent neural networks. In 2013 IEEE international conference on acoustics, speech and signal processing. IEEE, pp. 6645\u20136649","DOI":"10.1109\/ICASSP.2013.6638947"},{"issue":"1","key":"11388_CR42","doi-asserted-by":"publisher","first-page":"87","DOI":"10.1109\/TPAMI.2022.3152247","volume":"45","author":"K Han","year":"2022","unstructured":"Han K, Wang Y, Chen H, Chen X, Guo J, Liu Z, Tang Y, Xiao A, Xu C, Xu Y et al (2022) A survey on vision transformer. IEEE Trans Pattern Anal Mach Intell 45(1):87\u2013110","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11388_CR43","doi-asserted-by":"crossref","unstructured":"Hara K, Kataoka H, Satoh Y (2018) Can spatiotemporal 3D CNNs retrace the history of 2D CNNs and ImageNet? In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6546\u20136555","DOI":"10.1109\/CVPR.2018.00685"},{"key":"11388_CR44","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"issue":"8","key":"11388_CR45","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter S, Schmidhuber J (1997) Long short-term memory. Neural Comput 9(8):1735\u20131780","journal-title":"Neural Comput"},{"key":"11388_CR46","unstructured":"Howard AG, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Andreetto M, Adam H (2017) Mobilenets: efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861"},{"issue":"10","key":"11388_CR47","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0206049","volume":"13","author":"Y Hu","year":"2018","unstructured":"Hu Y, Wong Y, Wei W, Du Y, Kankanhalli M, Geng W (2018) A novel attention-based hybrid cnn-rnn architecture for semg-based gesture recognition. PLoS ONE 13(10):e0206049","journal-title":"PLoS ONE"},{"issue":"1","key":"11388_CR48","first-page":"3454167","volume":"2022","author":"A Hussain","year":"2022","unstructured":"Hussain A, Hussain T, Ullah W, Baik SW (2022) Vision transformer and deep sequence learning for human activity recognition in surveillance videos. Comput Intell Neurosci 2022(1):3454167","journal-title":"Comput Intell Neurosci"},{"key":"11388_CR49","doi-asserted-by":"publisher","first-page":"569","DOI":"10.1016\/j.aej.2023.05.050","volume":"74","author":"A Hussain","year":"2023","unstructured":"Hussain A, Khan SU, Khan N, Rida I, Alharbi M, Baik SW (2023) Low-light aware framework for human activity recognition via optimized dual stream parallel network. Alex Eng J 74:569\u2013583","journal-title":"Alex Eng J"},{"key":"11388_CR50","doi-asserted-by":"publisher","DOI":"10.1016\/j.engappai.2023.107218","volume":"127","author":"A Hussain","year":"2024","unstructured":"Hussain A, Khan SU, Khan N, Shabaz M, Baik SW (2024a) Ai-driven behavior biometrics framework for robust human activity recognition in surveillance systems. Eng Appl Artif Intell 127:107218","journal-title":"Eng Appl Artif Intell"},{"key":"11388_CR51","doi-asserted-by":"publisher","first-page":"632","DOI":"10.1016\/j.aej.2023.11.017","volume":"91","author":"A Hussain","year":"2024","unstructured":"Hussain A, Khan SU, Khan N, Ullah W, Alkhayyat A, Alharbi M, Baik SW (2024b) Shots segmentation-based optimized dual-stream framework for robust human activity recognition in surveillance video. Alex Eng J 91:632\u2013647","journal-title":"Alex Eng J"},{"key":"11388_CR52","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2023.102211","volume":"106","author":"A Hussain","year":"2024","unstructured":"Hussain A, Khan SU, Rida I, Khan N, Baik SW (2024c) Human centric attention with deep multiscale feature fusion framework for activity recognition in internet of medical things. Inf Fusion 106:102211","journal-title":"Inf Fusion"},{"key":"11388_CR53","doi-asserted-by":"crossref","unstructured":"Hussain A, Hussain T, Ullah W, Khan SU, Kim MJ, Muhammad K, Del\u00a0Ser J, Baik SW (2024d) Big data analysis for industrial activity recognition using attention-inspired sequential temporal convolution network. IEEE Trans Big Data","DOI":"10.1109\/TBDATA.2024.3489414"},{"key":"11388_CR54","doi-asserted-by":"crossref","unstructured":"Hussain A, Khan SU, Khan N, Bhatt MW, Farouk A, Bhola J, Baik SW (2024e) A hybrid transformer framework for efficient activity recognition using consumer electronics. IEEE Trans Consumer Electron","DOI":"10.1109\/TCE.2024.3373824"},{"key":"11388_CR55","doi-asserted-by":"crossref","unstructured":"Hussain A, Khan N, Munsif M, Kim MJ, SW Baik (2024f) Medium scale benchmark for cricket excited actions understanding. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp. 3399\u20133409","DOI":"10.1109\/CVPRW63382.2024.00344"},{"issue":"4","key":"11388_CR56","doi-asserted-by":"publisher","first-page":"447","DOI":"10.1016\/j.jksuci.2019.09.004","volume":"32","author":"N Jaouedi","year":"2020","unstructured":"Jaouedi N, Boujnah N, Bouhlel MS (2020) A new hybrid deep learning model for human action recognition. J King Saud Univ-Comput Inf Sci 32(4):447\u2013453","journal-title":"J King Saud Univ-Comput Inf Sci"},{"issue":"1","key":"11388_CR57","doi-asserted-by":"publisher","first-page":"221","DOI":"10.1109\/TPAMI.2012.59","volume":"35","author":"S Ji","year":"2012","unstructured":"Ji S, Xu W, Yang M, Yu K (2012) 3d convolutional neural networks for human action recognition. IEEE Trans Pattern Anal Mach Intell 35(1):221\u2013231","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11388_CR58","unstructured":"Jordan M (1986) Serial order: a parallel distributed processing approach. Technical Report, June 1985\u2013March 1986. Technical report, California Univ., San Diego, La Jolla (USA). Inst. for Cognitive Science"},{"key":"11388_CR59","doi-asserted-by":"crossref","unstructured":"Kalchbrenner N, Blunsom P (2013) Recurrent continuous translation models. In Proceedings of the 2013 conference on empirical methods in natural language processing, pp. 1700\u20131709","DOI":"10.18653\/v1\/D13-1176"},{"key":"11388_CR60","doi-asserted-by":"crossref","unstructured":"Karpathy A, Toderici G, Shetty S, Leung T, Sukthankar R, Fei-Fei L (2014) Large-scale video classification with convolutional neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1725\u20131732","DOI":"10.1109\/CVPR.2014.223"},{"issue":"10s","key":"11388_CR61","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3505244","volume":"54","author":"S Khan","year":"2022","unstructured":"Khan S, Naseer M, Hayat M, Zamir SW, Khan FS, Shah M (2022) Transformers in vision: a survey. ACM Comput Surv (CSUR) 54(10s):1\u201341","journal-title":"ACM Comput Surv (CSUR)"},{"issue":"5","key":"11388_CR62","doi-asserted-by":"publisher","first-page":"1366","DOI":"10.1007\/s11263-022-01594-9","volume":"130","author":"Y Kong","year":"2022","unstructured":"Kong Y, Fu Y (2022) Human action recognition and prediction: a survey. Int J Comput Vis 130(5):1366\u20131401","journal-title":"Int J Comput Vis"},{"key":"11388_CR63","unstructured":"Koot R, Hennerbichler M, Lu H (2021) Evaluating transformers for lightweight action recognition. arXiv preprint arXiv:2111.09641"},{"key":"11388_CR64","doi-asserted-by":"crossref","unstructured":"Kopuklu O, Kose N, Gunduz A, Rigoll G (2019) Resource efficient 3D convolutional neural networks. In Proceedings of the IEEE\/CVF international conference on computer vision workshops","DOI":"10.1109\/ICCVW.2019.00240"},{"key":"11388_CR65","unstructured":"Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. Adv Neural Inf Process Syst 25"},{"key":"11388_CR66","doi-asserted-by":"crossref","unstructured":"Kuehne H, Jhuang H, Poggio E, Poggio T, Serre T (2011) HMDB: a large video database for human motion recognition. In 2011 International conference on computer vision. IEEE, pp. 2556\u20132563","DOI":"10.1109\/ICCV.2011.6126543"},{"issue":"11","key":"11388_CR67","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","volume":"86","author":"Y LeCun","year":"1998","unstructured":"LeCun Y, Bottou L, Bengio Y, Haffner P (1998) Gradient-based learning applied to document recognition. Proc IEEE 86(11):2278\u20132324","journal-title":"Proc IEEE"},{"issue":"7553","key":"11388_CR68","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y LeCun","year":"2015","unstructured":"LeCun Y, Bengio Y, Hinton G (2015) Deep learning. Nature 521(7553):436\u2013444","journal-title":"Nature"},{"key":"11388_CR69","doi-asserted-by":"crossref","unstructured":"Lee S, Kim HG, Choi DH, Kim HI, Ro YM (2021) Video prediction recalling long-term motion context via memory alignment learning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp. 3054\u20133063","DOI":"10.1109\/CVPR46437.2021.00307"},{"key":"11388_CR70","unstructured":"Leong MC, Zhang H, Tan HL, Li L, Lim JH (2022) Combined CNN transformer encoder for enhanced fine-grained human action recognition. arXiv preprint arXiv:2208.01897"},{"key":"11388_CR73","doi-asserted-by":"crossref","unstructured":"Li Y, Lan C, Xing J, Zeng W, Yuan C, Liu J (2016) Online human action detection using joint classification-regression recurrent neural networks. In Computer vision\u2013ECCV 2016: 14th European conference, Amsterdam, The Netherlands, October 11\u201314, 2016, Proceedings, Part VII 14. Springer, pp. 203\u2013220","DOI":"10.1007\/978-3-319-46478-7_13"},{"key":"11388_CR71","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1016\/j.cviu.2017.10.011","volume":"166","author":"Z Li","year":"2018","unstructured":"Li Z, Gavrilyuk K, Gavves E, Jain M, Snoek CG (2018) Videolstm convolves, attends and flows for action recognition. Comput Vis Image Underst 166:41\u201350","journal-title":"Comput Vis Image Underst"},{"key":"11388_CR72","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2019.107037","volume":"98","author":"J Li","year":"2020","unstructured":"Li J, Liu X, Zhang M, Wang D (2020) Spatio-temporal deformable 3d convnets with attention for action recognition. Pattern Recogn 98:107037","journal-title":"Pattern Recogn"},{"key":"11388_CR74","doi-asserted-by":"crossref","unstructured":"Lin T, Wang Y, Liu X, Qiu X (2022) A survey of transformers AI Open 3:111\u2013132","DOI":"10.1016\/j.aiopen.2022.10.001"},{"key":"11388_CR76","doi-asserted-by":"crossref","unstructured":"Liu X, Qi DL, Xiao HB (2020) Construction and evaluation of the human behavior recognition model in kinematics under deep learning. J Ambient Intell Human Comput, 1\u20139","DOI":"10.1007\/s12652-020-02335-x"},{"key":"11388_CR75","doi-asserted-by":"crossref","unstructured":"Liu Z, Ning J, Cao Y, Wei Y, Zhang Z, Lin S, Hu H (2022) Video swim transformer. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp. 3202\u20133211","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"11388_CR77","doi-asserted-by":"crossref","unstructured":"Luong MT, Pham H, Manning CD (2015) Effective approaches to attention-based neural machine translation. arXiv preprint arXiv:1508.04025","DOI":"10.18653\/v1\/D15-1166"},{"issue":"1","key":"11388_CR78","doi-asserted-by":"publisher","first-page":"40","DOI":"10.3390\/signals4010002","volume":"4","author":"NUR Malik","year":"2023","unstructured":"Malik NUR, Abu-Bakar SAR, Sheikh UU, Channa A, Popescu N (2023) Cascading pose features with cnn-lstm for multiview human action recognition. Signals 4(1):40\u201355","journal-title":"Signals"},{"key":"11388_CR79","doi-asserted-by":"crossref","unstructured":"Mazzeo PL, Spagnolo P, Fasano M, Distante C (2022) Human action recognition with transformers. In International conference on image analysis and processing. Springer, pp. 230\u2013241","DOI":"10.1007\/978-3-031-06433-3_20"},{"key":"11388_CR80","unstructured":"Montgomery DC, Runger GC (2020) Applied statistics and probability for engineers. Wiley"},{"key":"11388_CR81","doi-asserted-by":"publisher","first-page":"820","DOI":"10.1016\/j.future.2021.06.045","volume":"125","author":"K Muhammad","year":"2021","unstructured":"Muhammad K, Ullah A, Imran AS, Sajjad M, Kiran MS, Sannino G, Albuquerque VHC et al (2021) Human action recognition using attention based lstm network with dilated cnn features. Future Gener Comput Syst 125:820\u2013830","journal-title":"Future Gener Comput Syst"},{"key":"11388_CR82","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2024.112480","volume":"304","author":"M Munsif","year":"2024","unstructured":"Munsif M, Khan SU, Khan N, Hussain A, Kim MJ, Baik SW (2024a) Contextual visual and motion salient fusion framework for action recognition in dark environments. Knowl-Based Syst 304:112480","journal-title":"Knowl-Based Syst"},{"key":"11388_CR83","doi-asserted-by":"crossref","unstructured":"Munsif M, Khan N, Hussain A, Kim MJ, Baik SW (2024b) Darkness-adaptive action recognition: leveraging efficient Tubelet slow-fast network for industrial applications. IEEE Trans Ind Inf","DOI":"10.1109\/TII.2024.3431070"},{"key":"11388_CR84","doi-asserted-by":"crossref","unstructured":"Mutegeki R, Han DS (2020) A CNN-LSTM approach to human activity recognition. In 2020 International conference on artificial intelligence in information and communication (ICAIIC). IEEE, pp. 362\u2013366","DOI":"10.1109\/ICAIIC48513.2020.9065078"},{"key":"11388_CR85","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2020.104078","volume":"106","author":"R Nayak","year":"2021","unstructured":"Nayak R, Pati UC, Das SK (2021) A comprehensive review on deep learning-based methods for video anomaly detection. Image Vis Comput 106:104078","journal-title":"Image Vis Comput"},{"key":"11388_CR86","doi-asserted-by":"publisher","first-page":"2259","DOI":"10.1007\/s10462-020-09904-8","volume":"54","author":"P Pareek","year":"2021","unstructured":"Pareek P, Thakkar A (2021) A survey on video-based human action recognition: recent updates, datasets, challenges, and applications. Artif Intell Rev 54:2259\u20132322","journal-title":"Artif Intell Rev"},{"issue":"3","key":"11388_CR87","doi-asserted-by":"publisher","first-page":"773","DOI":"10.1109\/TCSVT.2018.2808685","volume":"29","author":"Y Peng","year":"2018","unstructured":"Peng Y, Zhao Y, Zhang J (2018) Two-stream collaborative learning with spatial-temporal attention for video classification. IEEE Trans Circ Syst Video Technol 29(3):773\u2013786","journal-title":"IEEE Trans Circ Syst Video Technol"},{"key":"11388_CR88","doi-asserted-by":"crossref","unstructured":"Qiu Z, Yao T, Mei T (2017) Learning Spatio-temporal representation with pseudo-3D residual networks. In Proceedings of the IEEE international conference on computer vision, pp. 5533\u20135541","DOI":"10.1109\/ICCV.2017.590"},{"key":"11388_CR89","doi-asserted-by":"crossref","unstructured":"Reda DR, Chaieb F, Drira H, Aberkane A (2023) ConViViT-A deep neural network combining convolutions and factorized self-attention for human activity recognition. In 2023 IEEE 25th international workshop on multimedia signal processing (MMSP). IEEE, pp. 1\u20136","DOI":"10.1109\/MMSP59012.2023.10337696"},{"key":"11388_CR90","doi-asserted-by":"crossref","unstructured":"Rumelhart DE, Hinton GE, Williams RJ (1985) Learning internal representations by error propagation. Technical report, California Univ San Diego La Jolla Inst for Cognitive Science","DOI":"10.21236\/ADA164453"},{"issue":"5","key":"11388_CR91","doi-asserted-by":"publisher","first-page":"813","DOI":"10.1109\/TETCI.2020.3014367","volume":"5","author":"SP Sahoo","year":"2020","unstructured":"Sahoo SP, Ari S, Mahapatra K, Mohanty SP (2020) Har-depth: a novel framework for human action recognition using sequential learning and depth estimated history images. IEEE Trans Emerg Top Comput Intell 5(5):813\u2013825","journal-title":"IEEE Trans Emerg Top Comput Intell"},{"key":"11388_CR92","doi-asserted-by":"crossref","unstructured":"Schuldt C, Laptev I, Caputo B (2004) Recognizing human actions: a local SVM approach. In Proceedings of the 17th international conference on pattern recognition, 2004. ICPR 2004. IEEE, Volume\u00a03, pp. 32\u201336","DOI":"10.1109\/ICPR.2004.1334462"},{"key":"11388_CR93","doi-asserted-by":"crossref","unstructured":"Shan G, Ji Q, Xie Y (2021) Multi-view vision transformer for driver action recognition. In International conference on intelligent transportation engineering. Springer, pp. 970\u2013981","DOI":"10.1007\/978-981-19-2259-6_85"},{"key":"11388_CR94","unstructured":"Sharir G, Noy A, Zelnik-Manor L (2021) An image is worth 16x16 words, what is a video worth? arXiv preprint arXiv:2103.13915"},{"key":"11388_CR95","first-page":"6531","volume":"35","author":"S Shi","year":"2022","unstructured":"Shi S, Jiang L, Dai D, Schiele B (2022) Motion transformer with global intention localization and local movement refinement. Adv Neural Inf Process Syst 35:6531\u20136543","journal-title":"Adv Neural Inf Process Syst"},{"key":"11388_CR96","unstructured":"Simonyan K, Zisserman A (2014a) Two-stream convolutional networks for action recognition in videos. Adv Neural Inf Process Syst 27"},{"key":"11388_CR97","unstructured":"Simonyan K, Zisserman A (2014b) Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556"},{"key":"11388_CR98","unstructured":"Soomro K, Zamir AR, Shah M (2012) UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402"},{"key":"11388_CR99","unstructured":"Srivastava N, Mansimov E, Salakhudinov R (2015) Unsupervised learning of video representations using LSTMs. In International conference on machine learning. PMLR, pp. 843\u2013852"},{"issue":"3","key":"11388_CR100","first-page":"3200","volume":"45","author":"Z Sun","year":"2022","unstructured":"Sun Z, Ke Q, Rahmani H, Bennamoun M, Wang G, Liu J (2022) Human action recognition from various data modalities: a review. IEEE Trans Pattern Anal Mach Intell 45(3):3200\u20133225","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"issue":"14","key":"11388_CR101","doi-asserted-by":"publisher","first-page":"6384","DOI":"10.3390\/s23146384","volume":"23","author":"G Surek","year":"2023","unstructured":"Surek G, Seman L, Stefenon S, Mariani V, Coelho L (2023) Video-based human activity recognition using deep learning approaches. Sensors 23(14):6384","journal-title":"Sensors"},{"key":"11388_CR102","unstructured":"Sutskever I, Vinyals O, Le QV (2014) Sequence to sequence learning with neural networks. Adv Neural Inf Process Syst 27"},{"key":"11388_CR103","doi-asserted-by":"crossref","unstructured":"Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A (2015) Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1\u20139","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"11388_CR104","unstructured":"Touvron H, Cord M, Douze M, Massa F, Sablayrolles A, J\u00e9gou H (2021) Training data-efficient image transformers & distillation through attention. In International conference on machine learning. PMLR, pp. 10347\u201310357"},{"key":"11388_CR105","doi-asserted-by":"crossref","unstructured":"Tran D, Bourdev L, Fergus R, Torresani L, Paluri M (2015) Learning spatiotemporal features with 3D convolutional networks. In Proceedings of the IEEE international conference on computer vision, pp. 4489\u20134497","DOI":"10.1109\/ICCV.2015.510"},{"key":"11388_CR107","doi-asserted-by":"crossref","unstructured":"Tran D, Wang H, Torresani L, Ray J, LeCun Y, Paluri M (2018) A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6450\u20136459","DOI":"10.1109\/CVPR.2018.00675"},{"key":"11388_CR106","doi-asserted-by":"crossref","unstructured":"Tran D, Wang H, Torresani L, Feiszli M (2019) Video classification with channel-separated convolutional networks. In Proceedings of the IEEE\/CVF international conference on computer vision, pp. 5552\u20135561","DOI":"10.1109\/ICCV.2019.00565"},{"key":"11388_CR109","doi-asserted-by":"publisher","first-page":"1155","DOI":"10.1109\/ACCESS.2017.2778011","volume":"6","author":"A Ullah","year":"2017","unstructured":"Ullah A, Ahmad J, Muhammad K, Sajjad M, Baik SW (2017) Action recognition in video sequences using deep bi-directional lstm with cnn features. IEEE Access 6:1155\u20131166","journal-title":"IEEE Access"},{"key":"11388_CR108","unstructured":"Ulhaq A, Akhtar N, Pogrebna G, Mian A (2022) Vision transformers for action recognition: a survey. arXiv preprint arXiv:2209.05700"},{"issue":"6","key":"11388_CR110","doi-asserted-by":"publisher","first-page":"1510","DOI":"10.1109\/TPAMI.2017.2712608","volume":"40","author":"G Varol","year":"2017","unstructured":"Varol G, Laptev I, Schmid C (2017) Long-term temporal convolutions for action recognition. IEEE Trans Pattern Anal Mach Intell 40(6):1510\u20131517","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11388_CR111","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. Adv Neural Inf Process Syst 30"},{"key":"11388_CR116","doi-asserted-by":"crossref","unstructured":"Wang H, Schmid C (2013) Action recognition with improved trajectories. In Proceedings of the IEEE international conference on computer vision, pp. 3551\u20133558","DOI":"10.1109\/ICCV.2013.441"},{"key":"11388_CR112","unstructured":"Wang L, Xiong Y, Wang Z, Qiao Y (2015) Towards good practices for very deep two-stream convnets. arXiv preprint arXiv:1507.02159"},{"key":"11388_CR119","doi-asserted-by":"crossref","unstructured":"Wang L, Xiong Y, Wang Z, Qiao Y, Lin D, Tang X, Van\u00a0Gool L (2016) Temporal segment networks: towards good practices for deep action recognition. In European conference on computer vision. Springer, pp. 20\u201336","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"11388_CR115","doi-asserted-by":"crossref","unstructured":"Wang Y, Long M, Wang J, Yu PS (2017) Spatiotemporal pyramid network for video action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1529\u20131538","DOI":"10.1109\/CVPR.2017.226"},{"issue":"11","key":"11388_CR113","doi-asserted-by":"publisher","first-page":"2740","DOI":"10.1109\/TPAMI.2018.2868668","volume":"41","author":"L Wang","year":"2018","unstructured":"Wang L, Xiong Y, Wang Z, Qiao Y, Lin D, Tang X, Van Gool L (2018a) Temporal segment networks for action recognition in videos. IEEE Trans Pattern Anal Mach Intell 41(11):2740\u20132755","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11388_CR114","doi-asserted-by":"crossref","unstructured":"Wang X, Girshick R, Gupta A, He K (2018b) Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7794\u20137803","DOI":"10.1109\/CVPR.2018.00813"},{"key":"11388_CR117","doi-asserted-by":"crossref","unstructured":"Wang L, Tong Z, Ji B, Wu G (2021a) Tdn: Temporal difference networks for efficient action recognition. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp. 1895\u20131904","DOI":"10.1109\/CVPR46437.2021.00193"},{"key":"11388_CR118","unstructured":"Wang M, Xing J, Liu Y (2021b) Actionclip: a new paradigm for video action recognition. arXiv preprint arXiv:2109.08472"},{"key":"11388_CR120","doi-asserted-by":"crossref","unstructured":"Wu Z, Wang X, Jiang YG, Ye H, Xue X (2015) Modeling spatial-temporal clues in a hybrid deep learning framework for video classification. In Proceedings of the 23rd ACM international conference on multimedia, pp. 461\u2013470","DOI":"10.1145\/2733373.2806222"},{"key":"11388_CR121","doi-asserted-by":"publisher","first-page":"56855","DOI":"10.1109\/ACCESS.2020.2982225","volume":"8","author":"K Xia","year":"2020","unstructured":"Xia K, Huang J, Wang H (2020) Lstm-cnn architecture for human activity recognition. IEEE Access 8:56855\u201356866","journal-title":"IEEE Access"},{"key":"11388_CR122","doi-asserted-by":"crossref","unstructured":"Xie S, Sun C, Huang J, Tu Z, Murphy K (2018) Rethinking Spatiotemporal feature learning: speed-accuracy trade-offs in video classification. In Proceedings of the European conference on computer vision (ECCV), pp. 305\u2013321","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"11388_CR123","doi-asserted-by":"crossref","unstructured":"Xing Z, Dai Q, Hu H, Chen J, Wu Z, Jiang YG (2023) Svformer: semi-supervised video transformer for action recognition. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp. 18816\u201318826","DOI":"10.1109\/CVPR52729.2023.01804"},{"key":"11388_CR124","doi-asserted-by":"crossref","unstructured":"Ye X, Bilodeau GA (2023) A unified model for continuous conditional video prediction. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp. 3604\u20133613","DOI":"10.1109\/CVPRW59228.2023.00368"},{"key":"11388_CR125","doi-asserted-by":"crossref","unstructured":"Yin R, Yin J (2024) A two-stream hybrid CNN-transformer network for skeleton-based human interaction recognition. In Chinese conference on pattern recognition and computer vision (PRCV). Springer, pp. 395\u2013408","DOI":"10.1007\/978-981-97-8511-7_28"},{"key":"11388_CR126","doi-asserted-by":"crossref","unstructured":"Yue-Hei\u00a0Ng J, Hausknecht M, Vijayanarasimhan S, Vinyals O, Monga R, Toderici G (2015) Beyond short snippets: deep networks for video classification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4694\u20134702","DOI":"10.1109\/CVPR.2015.7299101"},{"key":"11388_CR127","doi-asserted-by":"crossref","unstructured":"Zhu L, Yang Y (2018) Compound memory networks for few-shot video classification. In Proceedings of the European conference on computer vision (ECCV), pp. 751\u2013766","DOI":"10.1007\/978-3-030-01234-2_46"},{"key":"11388_CR128","doi-asserted-by":"crossref","unstructured":"Zolfaghari M, Singh K, Brox T (2018) Eco: efficient convolutional network for online video understanding. In Proceedings of the European conference on computer vision (ECCV), pp. 695\u2013712","DOI":"10.1007\/978-3-030-01216-8_43"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11388-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-025-11388-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-025-11388-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T02:02:46Z","timestamp":1764986566000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-025-11388-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,17]]},"references-count":128,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2025,12]]}},"alternative-id":["11388"],"URL":"https:\/\/doi.org\/10.1007\/s10462-025-11388-3","relation":{},"ISSN":["1573-7462"],"issn-type":[{"value":"1573-7462","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,17]]},"assertion":[{"value":"14 October 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 August 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 October 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"387"}}