{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,30]],"date-time":"2026-01-30T22:53:46Z","timestamp":1769813626046,"version":"3.49.0"},"reference-count":82,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T00:00:00Z","timestamp":1765238400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T00:00:00Z","timestamp":1765238400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Complex Intell. Syst."],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Continuous Sign Language Recognition (CSLR) is fundamental to bridging the communication gap between hearing-impaired individuals and the broader society. The primary challenge lies in effectively modeling the complex spatial-temporal dynamic features in sign language videos. Current approaches typically employ independent processing strategies for motion feature extraction and temporal modeling, which impedes the unified modeling of action continuity and semantic integrity in sign language sequences. To address these limitations, we propose the Motion-Temporal Calibration Network (MTCNet), a novel framework for continuous sign language recognition that integrates dynamic feature enhancement and temporal calibration. The framework consists of two key innovative modules. First, the Cross-Frame Motion Refinement (CFMR) module implements an inter-frame differential attention mechanism combined with residual learning strategies, enabling precise motion feature modeling and effective enhancement of dynamic information between adjacent frames. Second, the Temporal-Channel Adaptive Recalibration (TCAR) module utilizes adaptive convolution kernel design and a dual-branch feature extraction architecture, facilitating joint optimization in both temporal and channel dimensions. In experimental evaluations, our method demonstrates competitive performance on the widely-used PHOENIX-2014 and PHOENIX-2014-T datasets, achieving results comparable to leading unimodal approaches. Moreover, it achieves state-of-the-art performance on the Chinese Sign Language (CSL) dataset. Through comprehensive ablation studies and quantitative analysis, we validate the effectiveness of our proposed method in fine-grained dynamic feature modeling and long-term dependency capture while maintaining computational efficiency.<\/jats:p>","DOI":"10.1007\/s40747-025-02156-5","type":"journal-article","created":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T03:41:13Z","timestamp":1765251673000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Motion-temporal calibration network for continuous sign language recognition"],"prefix":"10.1007","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-4847-4157","authenticated-orcid":false,"given":"Hongguan","family":"Hu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5869-8013","authenticated-orcid":false,"given":"Jianjun","family":"Peng","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1977-4674","authenticated-orcid":false,"given":"Zhidong","family":"Xiao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Li","family":"Guo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yi","family":"Hu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Di","family":"Wu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,12,9]]},"reference":[{"key":"2156_CR1","doi-asserted-by":"publisher","first-page":"1311","DOI":"10.1007\/s11263-018-1121-3","volume":"126","author":"O KOLLER","year":"2018","unstructured":"Koller O, Zargaran S, Ney H et al (2018) Deep sign: enabling robust statistical continuous sign Language recognition via hybrid CNN-HMMs [J]. Int J Comput Vision 126:1311\u20131325","journal-title":"Int J Comput Vision"},{"key":"2156_CR2","doi-asserted-by":"crossref","unstructured":"Hu H, Zhao W, Zhou W et al (2021) SignBERT: Pre-training of hand-model-aware representation for sign language recognition; proceedings of the Proceedings of the IEEE\/CVF international conference on computer vision, F, [C]","DOI":"10.1109\/ICCV48922.2021.01090"},{"key":"2156_CR3","doi-asserted-by":"crossref","unstructured":"Laines D, Gonzalez-Mendoza M, Ochoa-Ruiz G et al (2023) Isolated sign language recognition based on tree structure skeleton images; proceedings of the Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, F, [C]","DOI":"10.1109\/CVPRW59228.2023.00033"},{"key":"2156_CR4","doi-asserted-by":"publisher","first-page":"4271","DOI":"10.1109\/TMM.2023.3321502","volume":"26","author":"K LIN","year":"2024","unstructured":"Lin K, Wang X (2024) SKIM: Skeleton-Based isolated sign Language recognition with part mixing [J]. IEEE Trans Multimedia 26:4271\u20134280","journal-title":"IEEE Trans Multimedia"},{"issue":"31","key":"2156_CR5","doi-asserted-by":"publisher","first-page":"22177","DOI":"10.1007\/s11042-020-08961-z","volume":"79","author":"N ALOYSIUS","year":"2020","unstructured":"Aloysius N (2020) Understanding vision-based continuous sign Language recognition [J]. Multimedia Tools Appl 79(31):22177\u201322209","journal-title":"Multimedia Tools Appl"},{"key":"2156_CR6","doi-asserted-by":"publisher","first-page":"109903","DOI":"10.1016\/j.patcog.2023.109903","volume":"145","author":"L HU","year":"2024","unstructured":"Hu L, Gao L, Liu Z et al (2024) Scalable frame resolution for efficient continuous sign Language recognition [J]. Pattern Recogn 145:109903","journal-title":"Pattern Recogn"},{"key":"2156_CR7","doi-asserted-by":"publisher","first-page":"123695","DOI":"10.1016\/j.eswa.2024.123695","volume":"249","author":"R RASTGOO","year":"2024","unstructured":"Rastgoo R, Kiani K (2024) Word separation in continuous sign Language using isolated signs and post-processing [J]. Expert Syst Appl 249:123695","journal-title":"Expert Syst Appl"},{"key":"2156_CR8","doi-asserted-by":"crossref","unstructured":"Channayanamath M, Math A, Peddigari V et al (2021) [C] Dynamic hand gesture recognition using 3d-convolutional neural network; proceedings of the Communication Software and Networks: Proceedings of INDIA 2019, F, Springer","DOI":"10.1007\/978-981-15-5397-4_16"},{"key":"2156_CR9","doi-asserted-by":"crossref","unstructured":"Cui R, Liu H, Zhang C (2017) Recurrent convolutional neural networks for continuous sign language recognition by staged optimization; proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR.2017.175"},{"key":"#cr-split#-2156_CR10.1","doi-asserted-by":"crossref","unstructured":"Liang W, Xu X (2021) [C] Skeleton-based sign language recognition with attention-enhanced graph convolutional networks","DOI":"10.1007\/978-3-030-88480-2_62"},{"key":"#cr-split#-2156_CR10.2","unstructured":"proceedings of the Natural Language Processing and Chinese Computing: 10th CCF International Conference, NLPCC 2021, Qingdao, China, October 13-17, 2021, Proceedings, Part I 10, F, Springer"},{"key":"#cr-split#-2156_CR11.1","doi-asserted-by":"crossref","unstructured":"Camgoz Nc, Koller O et al (2020) [C] HADFIELD S,. Multi-channel transformers for multi-articulatory sign language translation","DOI":"10.1007\/978-3-030-66823-5_18"},{"key":"#cr-split#-2156_CR11.2","unstructured":"proceedings of the Computer Vision-ECCV 2020 Workshops: Glasgow, UK, August 23-28, 2020, Proceedings, Part IV 16, F, Springer"},{"issue":"10","key":"2156_CR12","first-page":"2428","volume":"45","author":"M JINKAI","year":"2024","unstructured":"Jinkai M, Jianjun P, Zhidong X et al (2024) Review of modular continuous sign Language recognition methods and techniques [J]. J Chin Comput Syst 45(10):2428\u20132441","journal-title":"J Chin Comput Syst"},{"key":"2156_CR13","doi-asserted-by":"crossref","unstructured":"Li D, Yu X, Xu C et al (2020) Transferring cross-domain knowledge for video sign language recognition; proceedings of the Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, F, [C]","DOI":"10.1109\/CVPR42600.2020.00624"},{"key":"2156_CR14","doi-asserted-by":"crossref","unstructured":"Pu J, Zhou W, Li H (2019) Iterative alignment network for continuous sign language recognition; proceedings of the Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR.2019.00429"},{"key":"2156_CR15","doi-asserted-by":"publisher","first-page":"113794","DOI":"10.1016\/j.eswa.2020.113794","volume":"164","author":"R RASTGOO","year":"2021","unstructured":"Rastgoo R, Kiani K, Escalera S (2021) Sign Language recognition: A deep survey [J]. Expert Syst Appl 164:113794","journal-title":"Expert Syst Appl"},{"key":"2156_CR16","doi-asserted-by":"crossref","unstructured":"Koller O (2017) ZARGARAN S, NEY H. Re-sign: Re-aligned end-to-end sequence modelling with deep recurrent CNN-HMMs; proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR.2017.364"},{"key":"2156_CR17","doi-asserted-by":"crossref","unstructured":"Carreira J (2017) ZISSERMAN A. Quo vadis, action recognition? a new model and the kinetics dataset; proceedings of the proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, F, [C]","DOI":"10.1109\/CVPR.2017.502"},{"key":"2156_CR18","doi-asserted-by":"crossref","unstructured":"Tran D, Wang H et al (2018) TORRESANI L,. A closer look at spatiotemporal convolutions for action recognition; proceedings of the Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, F, [C]","DOI":"10.1109\/CVPR.2018.00675"},{"key":"2156_CR19","doi-asserted-by":"crossref","unstructured":"Liu Z, Luo D, Wang Y et al (2020) Teinet: Towards an efficient architecture for video recognition; proceedings of the Proceedings of the AAAI conference on artificial intelligence, F, [C]","DOI":"10.1609\/aaai.v34i07.6836"},{"issue":"4","key":"2156_CR20","doi-asserted-by":"publisher","first-page":"4645","DOI":"10.1007\/s40747-023-00977-w","volume":"9","author":"Z CUI","year":"2023","unstructured":"Cui Z, Zhang W, Li Z et al (2023) Spatial\u2013temporal transformer for end-to-end sign Language recognition [J]. Complex Intell Syst 9(4):4645\u20134656","journal-title":"Complex Intell Syst"},{"key":"#cr-split#-2156_CR21.1","doi-asserted-by":"crossref","unstructured":"Saunders B, Camgoz N C Bowdenr (2020) Progressive transformers for end-to-end sign language production","DOI":"10.1007\/978-3-030-58621-8_40"},{"key":"#cr-split#-2156_CR21.2","unstructured":"proceedings of the Computer Vision-ECCV. : 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XI 16, F, 2020 [C]. Springer"},{"key":"2156_CR22","unstructured":"Vaswani A (2017) Attention is all you need [J]. Advances in Neural Information Processing Systems"},{"key":"2156_CR23","doi-asserted-by":"crossref","unstructured":"Lea C, Flynn M D, Vidal R et al (2017) Temporal convolutional networks for action segmentation and detection; proceedings of the proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, F, [C]","DOI":"10.1109\/CVPR.2017.113"},{"key":"2156_CR24","doi-asserted-by":"crossref","unstructured":"Wu C-Y, Li Y, Mangalam K et al (2022) Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition; proceedings of the Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, F, [C]","DOI":"10.1109\/CVPR52688.2022.01322"},{"key":"2156_CR25","doi-asserted-by":"publisher","first-page":"108","DOI":"10.1016\/j.cviu.2015.09.013","volume":"141","author":"O KOLLER","year":"2015","unstructured":"Koller O, Forster J (2015) Continuous sign Language recognition: towards large vocabulary statistical recognition systems handling multiple signers [J]. Comput Vis Image Underst 141:108\u2013125","journal-title":"Comput Vis Image Underst"},{"issue":"9","key":"2156_CR26","doi-asserted-by":"publisher","first-page":"2306","DOI":"10.1109\/TPAMI.2019.2911077","volume":"42","author":"H KOLLER O, CAMGOZ N C, NEY","year":"2019","unstructured":"Koller O, Camgoz N C, Ney H et al (2019) Weakly supervised learning with multi-stream CNN-LSTM-HMMs to discover sequential parallelism in sign Language videos [J]. IEEE Trans Pattern Anal Mach Intell 42(9):2306\u20132320","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"2156_CR27","doi-asserted-by":"crossref","unstructured":"PU J, ZHOU W, Hu H et al (2020) Boosting continuous sign language recognition via cross modality augmentation; proceedings of the Proceedings of the 28th ACM international conference on multimedia, F, [C]","DOI":"10.1145\/3394171.3413931"},{"key":"2156_CR28","unstructured":"Du Y, Chen Z, Xie H et al (2024) SVTRv2: CTC beats Encoder-Decoder models in scene text recognition [J]. arXiv preprint arXiv:241115858"},{"key":"2156_CR29","doi-asserted-by":"crossref","unstructured":"Fan R, Chu W, Chang P et al (2023) A ctc alignment-based non-autoregressive transformer for end-to-end automatic speech recognition [J]. IEEE\/ACM Transactions on Audio, Speech, and Language Processing, 31: 1436\u20131448","DOI":"10.1109\/TASLP.2023.3263789"},{"key":"2156_CR30","doi-asserted-by":"crossref","unstructured":"Min Y, Hao A, Chai X et al (2021) Visual alignment constraint for continuous sign language recognition; proceedings of the Proceedings of the IEEE\/CVF international conference on computer vision, F, [C]","DOI":"10.1109\/ICCV48922.2021.01134"},{"key":"2156_CR31","doi-asserted-by":"crossref","unstructured":"Hu L, Gao L, Liu Z et al (2023) Continuous sign language recognition with correlation network; proceedings of the Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, F, [C]","DOI":"10.1109\/CVPR52729.2023.00249"},{"key":"2156_CR32","doi-asserted-by":"crossref","unstructured":"Zheng J, Wang Y, Tan C et al (2023) Cvt-slr: Contrastive visual-textual transformation for sign language recognition with variational alignment; proceedings of the Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR52729.2023.02216"},{"key":"2156_CR33","doi-asserted-by":"publisher","first-page":"106587","DOI":"10.1016\/j.neunet.2024.106587","volume":"179","author":"L GAO","year":"2024","unstructured":"Gao L, Shi P (2024) Cross-modal knowledge distillation for continuous sign Language recognition [J]. Neural Netw 179:106587","journal-title":"Neural Netw"},{"key":"2156_CR34","doi-asserted-by":"crossref","unstructured":"Zhou H, Zhou W, Zhou Y et al (2020) Spatial-temporal multi-cue network for continuous sign language recognition; proceedings of the Proceedings of the AAAI conference on artificial intelligence, F, [C]","DOI":"10.1609\/aaai.v34i07.7001"},{"key":"2156_CR35","doi-asserted-by":"crossref","unstructured":"Zuo R, Mak B (2022) C2slr: Consistency-enhanced continuous sign language recognition; proceedings of the Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, F, [C]","DOI":"10.1109\/CVPR52688.2022.00507"},{"key":"2156_CR36","first-page":"17043","volume":"35","author":"Y CHEN","year":"2022","unstructured":"Chen Y, Zuo R, Wei F et al (2022) Two-stream network for sign Language recognition and translation [J]. Adv Neural Inf Process Syst 35:17043\u201317056","journal-title":"Adv Neural Inf Process Syst"},{"issue":"3","key":"2156_CR37","doi-asserted-by":"publisher","first-page":"500","DOI":"10.1109\/TPAMI.2010.143","volume":"33","author":"T BROX","year":"2010","unstructured":"Brox T, Malik J (2010) Large displacement optical flow: descriptor matching in variational motion Estimation [J]. IEEE Trans Pattern Anal Mach Intell 33(3):500\u2013513","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"2156_CR38","doi-asserted-by":"publisher","first-page":"123258","DOI":"10.1016\/j.eswa.2024.123258","volume":"248","author":"SB ABDULLAHI","year":"2024","unstructured":"Abdullahi Sb, Bolon-Canedo V Chamnongthaik et al (2024) Spatial\u2013temporal feature-based End-to-end fourier network for 3D sign Language recognition [J]. Expert Syst Appl 248:123258","journal-title":"Expert Syst Appl"},{"key":"2156_CR39","doi-asserted-by":"crossref","unstructured":"Wang X, Girshick R, Gupta A et al (2018) Non-local neural networks; proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR.2018.00813"},{"key":"2156_CR40","doi-asserted-by":"crossref","unstructured":"Wang H, Tran D, Torresani L et al (2020) Video modeling with correlation networks; proceedings of the Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR42600.2020.00043"},{"key":"#cr-split#-2156_CR41.1","doi-asserted-by":"crossref","unstructured":"Kwon H, Kim M, Kwak S et al (2020) Motionsqueeze: Neural motion feature learning for video understanding","DOI":"10.1007\/978-3-030-58517-4_21"},{"key":"#cr-split#-2156_CR41.2","unstructured":"proceedings of the Computer Vision-ECCV. : 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XVI 16, F, 2020 [C]. Springer"},{"key":"2156_CR42","doi-asserted-by":"crossref","unstructured":"Jiang B, Wang M, Gan W et al (2019) Stm: Spatiotemporal and motion encoding for action recognition; proceedings of the Proceedings of the IEEE\/CVF international conference on computer vision, F, [C]","DOI":"10.1109\/ICCV.2019.00209"},{"key":"2156_CR43","doi-asserted-by":"crossref","unstructured":"Li Y, Ji B, Shi X et al (2020) Tea: Temporal excitation and aggregation for action recognition; proceedings of the Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR42600.2020.00099"},{"key":"2156_CR44","doi-asserted-by":"crossref","unstructured":"Wang Z, She Q, Smolic A (2021) Action-net: Multipath excitation for action recognition; proceedings of the Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR46437.2021.01301"},{"key":"2156_CR45","unstructured":"Huang Z, Xu W, Yu K (2015) Bidirectional LSTM-CRF models for sequence tagging [J]. arXiv preprint arXiv:150801991"},{"key":"2156_CR46","doi-asserted-by":"crossref","unstructured":"Wang L, Tong Z, Ji B et al (2021) Tdn: Temporal difference networks for efficient action recognition; proceedings of the Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR46437.2021.00193"},{"key":"2156_CR47","doi-asserted-by":"crossref","unstructured":"Liu Z, Ning J, Cao Y et al (2022) Video swin transformer; proceedings of the Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"2156_CR48","doi-asserted-by":"crossref","unstructured":"Liu Z, Lin Y, Cao Y et al (2021) Swin transformer: Hierarchical vision transformer using shifted windows; proceedings of the Proceedings of the IEEE\/CVF international conference on computer vision, F, [C]","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"2156_CR49","unstructured":"Zhang S, Guo S, Huang W et al (2020) V4d: 4d convolutional neural networks for video-level representation learning [J]. arXiv preprint arXiv:200207442"},{"key":"2156_CR50","unstructured":"Bai S, Kolter J Z Koltunv (2018) An empirical evaluation of generic convolutional and recurrent networks for sequence modeling [J]. arXiv preprint arXiv:180301271"},{"key":"2156_CR51","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S et al (2016) Deep residual learning for image recognition; proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR.2016.90"},{"key":"2156_CR52","doi-asserted-by":"crossref","unstructured":"Camgoz N C, Hadfield S, Koller O et al (2018) Neural sign language translation; proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR.2018.00812"},{"key":"2156_CR53","doi-asserted-by":"crossref","unstructured":"Huang J, Zhou W, Zhang Q et al (2018) Video-based sign language recognition without temporal segmentation; proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, F, [C]","DOI":"10.1609\/aaai.v32i1.11903"},{"key":"2156_CR54","unstructured":"Loshchilov I (2017) Decoupled weight decay regularization [J]. arXiv preprint arXiv:171105101"},{"key":"2156_CR55","unstructured":"Loshchilov I, Hutter F (2016) Sgdr Stochastic gradient descent with warm restarts [J]. arXiv preprint arXiv:160803983"},{"key":"#cr-split#-2156_CR56.1","doi-asserted-by":"crossref","unstructured":"Niu Z, Mak B (2020) Stochastic fine-grained labeling of multi-state sign glosses for continuous sign language recognition","DOI":"10.1007\/978-3-030-58517-4_11"},{"key":"#cr-split#-2156_CR56.2","unstructured":"proceedings of the Computer Vision-ECCV : 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XVI 16, F, 2020 [C]. Springer"},{"key":"2156_CR57","unstructured":"Camgoz N C, Koller O, Hadfield S et al (2020) Sign language transformers: Joint end-to-end sign language recognition and translation; proceedings of the Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, F, [C]"},{"key":"#cr-split#-2156_CR58.1","doi-asserted-by":"crossref","unstructured":"Cheng Kl, Yang Z, Chen Q et al (2020) Fully convolutional networks for continuous sign language recognition","DOI":"10.1007\/978-3-030-58586-0_41"},{"key":"#cr-split#-2156_CR58.2","unstructured":"proceedings of the Computer Vision-ECCV. : 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XXIV 16, F, 2020 [C]. Springer"},{"key":"2156_CR59","doi-asserted-by":"crossref","unstructured":"Zhou H, Zhou W, Qi W et al (2021) Improving sign language translation with monolingual data by sign back-translation; proceedings of the Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, F, [C]","DOI":"10.1109\/CVPR46437.2021.00137"},{"key":"2156_CR60","doi-asserted-by":"crossref","unstructured":"Chen Y, Wei F, Sun X et al (2022) A simple multi-modality transfer learning baseline for sign language translation; proceedings of the Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, F, [C]","DOI":"10.1109\/CVPR52688.2022.00506"},{"key":"2156_CR61","first-page":"4810","volume":"2022","author":"R ZUO","year":"2022","unstructured":"Zuo R (2022) Local Context-aware Self-attention for continuous sign Language Recognition}} [J]. Proc Interspeech 2022:4810\u20134814","journal-title":"Proc Interspeech"},{"key":"2156_CR62","doi-asserted-by":"crossref","unstructured":"Hao A, Min Y, Chen X (2021) Self-mutual distillation learning for continuous sign language recognition; proceedings of the Proceedings of the IEEE\/CVF international conference on computer vision, F, [C]","DOI":"10.1109\/ICCV48922.2021.01111"},{"key":"2156_CR63","doi-asserted-by":"crossref","unstructured":"Jang Y, Oh Y, Cho J W et al (2023) [C] Self-sufficient framework for continuous sign language recognition; proceedings of the ICASSP 2023\u20132023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), F, IEEE","DOI":"10.1109\/ICASSP49357.2023.10095732"},{"key":"2156_CR64","doi-asserted-by":"crossref","unstructured":"Hu L, Gao L, Liu Z et al (2022) [C] Temporal lift pooling for continuous sign language recognition; proceedings of the European conference on computer vision, F, Springer","DOI":"10.1007\/978-3-031-19833-5_30"},{"key":"2156_CR65","doi-asserted-by":"crossref","unstructured":"Hu L, Gao L, Liu Z et al (2023) Self-emphasizing network for continuous sign language recognition; proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, F, [C]","DOI":"10.1609\/aaai.v37i1.25164"},{"key":"2156_CR66","doi-asserted-by":"crossref","unstructured":"Senhua X, Liqing G, Liang W et al (2024) Multi-scale context-aware network for continuous sign Language recognition [J], vol 6. Virtual Reality & Intelligent Hardware, pp 323\u2013337. 4","DOI":"10.1016\/j.vrih.2023.06.011"},{"key":"2156_CR67","doi-asserted-by":"crossref","unstructured":"Wang Z, Li D, Jiang R et al (2025) Continuous sign Language recognition with Multi-Scale Spatial-Temporal feature enhancement [J]. IEEE Access","DOI":"10.1109\/ACCESS.2025.3526330"},{"key":"2156_CR68","doi-asserted-by":"crossref","unstructured":"Aloysius N, Geetha M, Nedungadi P (2025) Optimized Multi-Modal Conformer-based framework for continuous sign Language recognition [J]. IEEE Open J Comput Soc","DOI":"10.1109\/OJCS.2025.3564828"},{"key":"2156_CR69","unstructured":"Aloysius N, Nedungadi P (2024) Continuous Sign Language Recognition with Adapted Conformer via Unsupervised Pretraining [J]. arXiv preprint arXiv:240512018"},{"key":"2156_CR70","unstructured":"Chen H, Wang J, Guo Z et al (2024) SignVTCL: Multi-Modal continuous sign Language recognition enhanced by Visual-Textual contrastive learning [J]. arXiv preprint arXiv:240111847"},{"key":"2156_CR71","doi-asserted-by":"crossref","unstructured":"Cihan Camgoz N, Hadfield S, Koller O et al (2017) Subunets: End-to-end hand shape and continuous sign language recognition; proceedings of the Proceedings of the IEEE international conference on computer vision, F, [C]","DOI":"10.1109\/ICCV.2017.332"},{"key":"2156_CR72","unstructured":"Yang Z, Shi Z, Shen X et al (2019) Sf-net: structured feature network for continuous sign Language recognition [J]. arXiv preprint arXiv:190801341"},{"key":"2156_CR73","unstructured":"Zhu Q, Li J, Yuan F et al (2024) Continuous sign Language recognition based on cross-resolution knowledge distillation [J]. Arab J Sci Eng, : 1\u201313"},{"key":"2156_CR74","unstructured":"Hu L, Shi T, GAO L et al (2024) Improving continuous sign Language recognition with adapted image models [J]. arXiv preprint arXiv:240408226"},{"issue":"2","key":"2156_CR75","doi-asserted-by":"publisher","first-page":"023059","DOI":"10.1117\/1.JEI.33.2.023059","volume":"33","author":"Q ZHU","year":"2024","unstructured":"Zhu Q, Li J, Yuan F et al (2024) Multiscale Temporal network for continuous sign Language recognition [J]. J Electron Imaging 33(2):023059\u2013023059","journal-title":"J Electron Imaging"},{"issue":"6","key":"2156_CR76","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3640815","volume":"20","author":"R ZUO","year":"2024","unstructured":"Zuo R, Mak B (2024) Improving continuous sign Language recognition with consistency constraints and signer removal [J]. ACM Trans Multimedia Comput Commun Appl 20(6):1\u201325","journal-title":"ACM Trans Multimedia Comput Commun Appl"}],"container-title":["Complex &amp; Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s40747-025-02156-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s40747-025-02156-5","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s40747-025-02156-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,30]],"date-time":"2026-01-30T11:48:51Z","timestamp":1769773731000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s40747-025-02156-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,9]]},"references-count":82,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["2156"],"URL":"https:\/\/doi.org\/10.1007\/s40747-025-02156-5","relation":{},"ISSN":["2199-4536","2198-6053"],"issn-type":[{"value":"2199-4536","type":"print"},{"value":"2198-6053","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,9]]},"assertion":[{"value":"23 June 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 October 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 December 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no relevant financial or non-financial interests to disclose in relation to this research. All authors certify that they have no affiliations with or involvement in any organization or entity with any financial or non-financial interest in the subject matter discussed in this manuscript. The authors declare that they have no competing interests that could have influenced the work reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"35"}}