{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:50:59Z","timestamp":1782402659045,"version":"3.54.5"},"reference-count":78,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100020831","name":"Australian Resuscitation Council","doi-asserted-by":"publisher","award":["DP220100800"],"award-info":[{"award-number":["DP220100800"]}],"id":[{"id":"10.13039\/501100020831","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100020831","name":"Australian Resuscitation Council","doi-asserted-by":"publisher","award":["DE230100477"],"award-info":[{"award-number":["DE230100477"]}],"id":[{"id":"10.13039\/501100020831","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Like spoken languages, a single sign language expression could correspond to multiple valid textual interpretations. Hence, learning a rigid one-to-one mapping for sign language translation (SLT) models might be inadequate, particularly in the case of limited data. In this work, we introduce a Diverse Sign Language Translation (DivSLT) task, aiming to generate diverse yet accurate translations for sign language videos. Firstly, we employ large language models (LLM) to generate multiple references for the widely-used CSL-Daily and PHOENIX14T SLT datasets. Here, native speakers are only invited to touch up inaccurate references, thus significantly improving the annotation efficiency. Secondly, we provide a benchmark model to spur research in this task. Specifically, we investigate multi-reference training strategies enabling our DivSLT model to achieve diverse translations. Then, to enhance translation accuracy, we employ the max-reward-driven reinforcement learning objective that maximizes the reward of the translated result. Additionally, we utilize multiple metrics to assess the accuracy, diversity, and semantic precision of the DivSLT task. Experimental results on the enriched datasets demonstrate that our DivSLT method achieves not only better translation performance but also diverse translation results.<\/jats:p>","DOI":"10.1007\/s11263-026-02900-5","type":"journal-article","created":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:41:09Z","timestamp":1782402069000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Diverse Sign Language Translation"],"prefix":"10.1007","volume":"134","author":[{"given":"Xin","family":"Shen","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lei","family":"Shen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shaozu","family":"Yuan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Heming","family":"Du","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haiyang","family":"Sun","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xin","family":"Yu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"2900_CR1","unstructured":"Albanie, S., Varol, G., Momeni, L., Bull, H., Afouras, T., Chowdhury, H., Fox, N., Woll, B., Cooper, R., McParland, A., & Zisserman, A. (2021). Bbc-oxford british sign language dataset. CoRR arXiv:2111.03635."},{"key":"2900_CR2","unstructured":"Bahdanau, D., Brakel, P., Xu, K., Goyal, A., Lowe, R., Pineau, J., Courville, A. C., & Bengio, Y. (2017). An actor-critic algorithm for sequence prediction. In: 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net. https:\/\/openreview.net\/forum?id=SJDaqqveg."},{"key":"2900_CR3","doi-asserted-by":"publisher","unstructured":"Bull, H., Afouras, T., Varol, G., Albanie, S., Momeni, L., & Zisserman, A. (2021). Aligning subtitles in sign language videos. In: 2021 IEEE\/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021, pp. 11532\u201311541. IEEE. https:\/\/doi.org\/10.1109\/ICCV48922.2021.01135 .","DOI":"10.1109\/ICCV48922.2021.01135"},{"key":"2900_CR4","doi-asserted-by":"publisher","unstructured":"Camg\u00f6z, N.C., Hadfield, S., Koller, O., Ney, H., & Bowden, R. (2018). Neural sign language translation. In: 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pp. 7784\u20137793. Computer Vision Foundation \/ IEEE Computer Society. https:\/\/doi.org\/10.1109\/CVPR.2018.00812 . http:\/\/openaccess.thecvf.com\/content_cvpr_2018\/html\/Camgoz_Neural_Sign_Language_CVPR_2018_paper.html.","DOI":"10.1109\/CVPR.2018.00812"},{"key":"2900_CR5","doi-asserted-by":"publisher","unstructured":"Camg\u00f6z, N. C., Koller, O., Hadfield, S., & Bowden, R. (2020). Sign language transformers: Joint end-to-end sign language recognition and translation. In: 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pp. 10020\u201310030. Computer Vision Foundation \/ IEEE. https:\/\/doi.org\/10.1109\/CVPR42600.2020.01004 . https:\/\/openaccess.thecvf.com\/content_CVPR_2020\/html\/Camgoz_Sign_Language_Transformers_Joint_End-to-End_Sign_Language_Recognition_and_Translation_CVPR_2020_paper.html","DOI":"10.1109\/CVPR42600.2020.01004"},{"key":"2900_CR6","doi-asserted-by":"publisher","unstructured":"Chen, Y., Wei, F., Sun, X., Wu, Z., & Lin, S. (2022). A simple multi-modality transfer learning baseline for sign language translation. In: IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pp. 5110\u20135120. IEEE. https:\/\/doi.org\/10.1109\/CVPR52688.2022.00506.","DOI":"10.1109\/CVPR52688.2022.00506"},{"key":"2900_CR7","doi-asserted-by":"crossref","unstructured":"Chen, Z., Zhou, B., Huang, Y., Wan, J., Hu, Y., Shi, H., Liang, Y., Lei, Z., & Zhang, D. (2025). C 2 rl: Content and context representation learning for gloss-free sign language translation and retrieval. IEEE Transactions on Circuits and Systems for Video Technology.","DOI":"10.1109\/TCSVT.2025.3553052"},{"key":"2900_CR8","doi-asserted-by":"crossref","unstructured":"Chen, Z., Zhou, B., Li, J., Wan, J., Lei, Z., Jiang, N., Lu, Q., & Zhao, G. (2024). Factorized learning assisted with large language model for gloss-free sign language translation. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 7071\u20137081).","DOI":"10.63317\/4qo2dsetvyk8"},{"key":"2900_CR9","doi-asserted-by":"crossref","unstructured":"Chen, Y., Zuo, R., Wei, F., Wu, Y., Liu, S., & Mak, B. (2022). Two-stream network for sign language recognition and translation. In: NeurIPS. http:\/\/papers.nips.cc\/paper_files\/paper\/2022\/hash\/6cd3ac24cdb789beeaa9f7145670fcae-Abstract-Conference.html","DOI":"10.52202\/068431-1240"},{"key":"2900_CR10","doi-asserted-by":"publisher","unstructured":"Deng, J., Dong, W., Socher, R., Li, L., Li, K., & Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. In: 2009 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2009), 20-25 June 2009, Miami, Florida, USA, pp. 248\u2013255. IEEE Computer Society. https:\/\/doi.org\/10.1109\/CVPR.2009.5206848 .","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"2900_CR11","doi-asserted-by":"crossref","unstructured":"Devlin, J., Chang, M., Lee, K., & Toutanova, K. (2019). BERT: pre-training of deep bidirectional transformers for language understanding. In: Burstein, J., Doran, C., Solorio, T. (eds.) Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics, pp. 4171\u20134186. Association for Computational Linguistics.","DOI":"10.18653\/v1\/N19-1423"},{"key":"2900_CR12","unstructured":"Dreyer, M., & Marcu, D. (2012). Hyter: Meaning-equivalent semantics for translation evaluation. In: Human Language Technologies: Conference of the North American Chapter of the Association of Computational Linguistics, Proceedings, June 3-8, Montr\u00e9al, Canada, pp. 162\u2013171. The Association for Computational Linguistics, (2012). https:\/\/aclanthology.org\/N12-1017\/"},{"key":"2900_CR13","doi-asserted-by":"publisher","unstructured":"Duarte, A.C., Palaskar, S., Ventura, L., Ghadiyaram, D., DeHaan, K., Metze, F., Torres, J., & Gir\u00f3-i-Nieto, X. (2021). How2sign: A large-scale multimodal dataset for continuous american sign language. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, Virtual, June 19-25, pp. 2735\u20132744. Computer Vision Foundation \/ IEEE, (2021). https:\/\/doi.org\/10.1109\/CVPR46437.2021.00276 . https:\/\/openaccess.thecvf.com\/content\/CVPR2021\/html\/Duarte_How2Sign_A_Large-Scale_Multimodal_Dataset_for_Continuous_American_Sign_Language_CVPR_2021_paper.html","DOI":"10.1109\/CVPR46437.2021.00276"},{"key":"2900_CR14","doi-asserted-by":"publisher","unstructured":"Edunov, S., Ott, M., Auli, M., Grangier, D., & Ranzato, M. (2018). Classical structured prediction losses for sequence to sequence learning. In: Walker, M.A., Ji, H., Stent, A. (eds.) Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pp. 355\u2013364. Association for Computational Linguistics. https:\/\/doi.org\/10.18653\/V1\/N18-1033 .","DOI":"10.18653\/V1\/N18-1033"},{"key":"2900_CR15","doi-asserted-by":"crossref","unstructured":"Gong, J., Foo, L. G., He, Y., Rahmani, H., & Liu, J. (2024). Llms are good sign language translators. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 18362\u201318372).","DOI":"10.1109\/CVPR52733.2024.01738"},{"key":"2900_CR16","unstructured":"Hanke, T., Schulder, M., Konrad, R., & Jahn, E. (2020). Extending the public dgs corpus in size and depth. Sign-lang@ LREC 2020 (pp. 75\u201382) European Language Resources Association (ELRA)."},{"key":"2900_CR17","unstructured":"He, D., Xia, Y., Qin, T., Wang, L., Yu, N., Liu, T., & Ma, W. (2016). Dual learning for machine translation. In: Lee, D.D., Sugiyama, M., Luxburg, U., Guyon, I., Garnett, R. (eds.) Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain, pp. 820\u2013828. https:\/\/proceedings.neurips.cc\/paper\/2016\/hash\/5b69b9cb83065d403869739ae7f0995e-Abstract.html."},{"key":"2900_CR18","doi-asserted-by":"publisher","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pp. 770\u2013778. IEEE Computer Society. https:\/\/doi.org\/10.1109\/CVPR.2016.90 .","DOI":"10.1109\/CVPR.2016.90"},{"key":"2900_CR19","doi-asserted-by":"publisher","unstructured":"Hu, H., Pu, J., Zhou, W., Fang, H., & Li, H. (2023). Prior-aware cross modality augmentation learning for continuous sign language recognition. IEEE Transactions on Multimedia, 1\u201314. https:\/\/doi.org\/10.1109\/TMM.2023.3268368.","DOI":"10.1109\/TMM.2023.3268368"},{"key":"2900_CR20","unstructured":"Jiao, W., Wang, W., Huang, J.-t., Wang, X., & Tu, Z. (2023). Is chatgpt a good translator? a preliminary study. arXiv preprint arXiv:2301.08745."},{"key":"2900_CR21","doi-asserted-by":"crossref","unstructured":"Jiao, P., Min, Y., & Chen, X. (2024). Visual alignment pre-training for sign language translation. European Conference on Computer Vision (pp. 349\u2013367). Springer.","DOI":"10.1007\/978-3-031-72946-1_20"},{"key":"2900_CR22","doi-asserted-by":"crossref","unstructured":"Karpathy, A., & Fei-Fei, L. (2015). Deep visual-semantic alignments for generating image descriptions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 3128\u20133137).","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"2900_CR23","doi-asserted-by":"publisher","unstructured":"Khayrallah, H., Thompson, B., Post, M., & Koehn, P. (2020). Simulated multiple reference training improves low-resource machine translation. In: Webber, B., Cohn, T., He, Y., Liu, Y. (eds.) Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pp. 82\u201389. Association for Computational Linguistics. https:\/\/doi.org\/10.18653\/V1\/2020.EMNLP-MAIN.7.","DOI":"10.18653\/V1\/2020.EMNLP-MAIN.7"},{"key":"2900_CR24","doi-asserted-by":"crossref","unstructured":"Ko, S., Kim, C. J., Jung, H., & Cho, C. S. (2018). Neural sign language translation based on human keypoint estimation. CoRR arXiv:1811.11436.","DOI":"10.3390\/app9132683"},{"key":"2900_CR25","doi-asserted-by":"publisher","unstructured":"Lachaux, M., Joulin, A., & Lample, G. (2020). Target conditioning for one-to-many generation. In: Cohn, T., He, Y., Liu, Y. (eds.) Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020. Findings of ACL, vol. EMNLP 2020, pp. 2853\u20132862. Association for Computational Linguistics. https:\/\/doi.org\/10.18653\/V1\/2020.FINDINGS-EMNLP.256 .","DOI":"10.18653\/V1\/2020.FINDINGS-EMNLP.256"},{"key":"2900_CR26","unstructured":"Li, J., Monroe, W., & Jurafsky, D. (2016). A simple, fast diverse decoding algorithm for neural generation arXiv preprint arXiv:1611.08562."},{"key":"2900_CR27","doi-asserted-by":"crossref","unstructured":"Li, D., Rodriguez, C., Yu, X., & Li, H. (2020). Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. In: The IEEE Winter Conference on Applications of Computer Vision, pp. 1459\u20131469.","DOI":"10.1109\/WACV45572.2020.9093512"},{"key":"2900_CR28","unstructured":"Li, D., Xu, C., Yu, X., Zhang, K., Swift, B., Suominen, H., & Li, H. (2020). Tspnet: Hierarchical feature learning via temporal semantic pyramid for sign language translation. In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H. (eds.) Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, Virtual. https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/8c00dee24c9878fea090ed070b44f1ab-Abstract.html."},{"key":"2900_CR29","unstructured":"Li, Z., Zhou, W., Zhao, W., Wu, K., Hu, H., & Li, H. (2025). Uni-sign: Toward unified sign language understanding at scale. In: The Thirteenth International Conference on Learning Representations."},{"key":"2900_CR30","unstructured":"Lin, C.-Y. (2004). Rouge: A package for automatic evaluation of summaries. Text Summarization Branches Out (pp. 74\u201381)."},{"key":"2900_CR31","doi-asserted-by":"crossref","unstructured":"Lin, K., Wang, X., Zhu, L., Sun, K., Zhang, B., & Yang, Y. (2023). Gloss-free end-to-end sign language translation. In: Rogers, A., Boyd-Graber, J.L., Okazaki, N. (eds.) Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pp. 12904\u201312916. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2023.acl-long.722"},{"key":"2900_CR32","doi-asserted-by":"publisher","unstructured":"Lin, H., Yang, B., Yao, L., Liu, D., Zhang, H., Xie, J., Zhang, M., & Su, J. (2022). Bridging the gap between training and inference: Multi-candidate optimization for diverse neural machine translation. In: Carpuat, M., Marneffe, M., Ru\u00edz, I.V.M. (eds.) Findings of the Association for Computational Linguistics: NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pp. 2622\u20132632. Association for Computational Linguistics. https:\/\/doi.org\/10.18653\/V1\/2022.FINDINGS-NAACL.200.","DOI":"10.18653\/V1\/2022.FINDINGS-NAACL.200"},{"key":"2900_CR33","doi-asserted-by":"crossref","unstructured":"Lin, H., Yao, L., Yang, B., Liu, D., Zhang, H., Luo, W., Huang, D., & Su, J. (2021). Towards user-driven neural machine translation. In C. Zong, F. Xia, W. Li, & R. Navigli (Eds.), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL\/IJCNLP.Association for Computational Linguistics.","DOI":"10.18653\/v1\/2021.acl-long.310"},{"key":"2900_CR34","doi-asserted-by":"crossref","unstructured":"Liu, Y., Han, T., Ma, S., Zhang, J., Yang, Y., Tian, J., He, H., Li, A., He, M., Liu, Z., and others. (2023). Summary of chatgpt-related research and perspective towards the future of large language models. Meta-Radiology, 100017.","DOI":"10.1016\/j.metrad.2023.100017"},{"key":"2900_CR35","doi-asserted-by":"publisher","first-page":"726","DOI":"10.1162\/TACL_A_00343","volume":"8","author":"Y Liu","year":"2020","unstructured":"Liu, Y., Gu, J., Goyal, N., Li, X., Edunov, S., Ghazvininejad, M., Lewis, M., & Zettlemoyer, L. (2020). Multilingual denoising pre-training for neural machine translation. Trans. Assoc. Comput. Linguistics, 8, 726\u2013742. https:\/\/doi.org\/10.1162\/TACL_A_00343","journal-title":"Trans. Assoc. Comput. Linguistics"},{"key":"2900_CR36","doi-asserted-by":"crossref","unstructured":"Mehrish, A., Majumder, N., Bharadwaj, R., Mihalcea, R., & Poria, S. (2023). A review of deep learning techniques for speech processing. Information Fusion, 101869.","DOI":"10.1016\/j.inffus.2023.101869"},{"key":"2900_CR37","doi-asserted-by":"publisher","unstructured":"Michel, P., & Neubig, G. (2018). Extreme adaptation for personalized neural machine translation. In: Gurevych, I., Miyao, Y. (eds.) Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 2: Short Papers, pp. 312\u2013318. Association for Computational Linguistics. https:\/\/doi.org\/10.18653\/V1\/P18-2050 . https:\/\/aclanthology.org\/P18-2050\/","DOI":"10.18653\/V1\/P18-2050"},{"key":"2900_CR38","unstructured":"Moryossef, A., Yin, K., Neubig, G., & Goldberg, Y. (2021). Data augmentation for sign language gloss translation. In: Shterionov, D. (ed.) Proceedings of the 1st International Workshop on Automatic Translation for Signed and Spoken Languages, pp. 1\u201311. Association for Machine Translation in the Americas. https:\/\/aclanthology.org\/2021.mtsummit-at4ssl.1"},{"key":"2900_CR39","unstructured":"OpenAI: ChatGPT. Available at https:\/\/www.openai.com\/chatgpt\/ (Accessed: 2023-10-30) (2023)"},{"key":"2900_CR40","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). Bleu: a method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311\u2013318).","DOI":"10.3115\/1073083.1073135"},{"key":"2900_CR41","doi-asserted-by":"crossref","unstructured":"Pasunuru, R., & Bansal, M. (2017). Reinforced video captioning with entailment rewards. In: Palmer, M., Hwa, R., Riedel, S. (eds.) Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017.","DOI":"10.18653\/v1\/D17-1103"},{"key":"2900_CR42","unstructured":"Paulus, R., Xiong, C., & Socher, R. (2018). A deep reinforced model for abstractive summarization. In: 6th International Conference on Learning Representations, ICLR 2018."},{"key":"2900_CR43","doi-asserted-by":"crossref","unstructured":"Pu, A., Chung, H. W., Parikh, A. P., Gehrmann, S., & Sellam, T. (2021). Learning compact metrics for mt. In: Proceedings of EMNLP.","DOI":"10.18653\/v1\/2021.emnlp-main.58"},{"key":"2900_CR44","unstructured":"Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In Meila, M., Zhang, T. (Eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021 (pp. 18\u201324). Virtual Event."},{"key":"2900_CR45","unstructured":"Ranzato, M., Chopra, S., Auli, M., & Zaremba, W. (2016). Sequence level training with recurrent neural networks. In: Bengio, Y., LeCun, Y. (eds.) 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings. arXiv:1511.06732."},{"key":"2900_CR46","doi-asserted-by":"publisher","unstructured":"Sellam, T., Das, D., & Parikh, A. P. (2020). BLEURT: learning robust metrics for text generation. In: Jurafsky, D., Chai, J., Schluter, N., Tetreault, J.R. (eds.) Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pp. 7881\u20137892. Association for Computational Linguistics. https:\/\/doi.org\/10.18653\/V1\/2020.ACL-MAIN.704.","DOI":"10.18653\/V1\/2020.ACL-MAIN.704"},{"key":"2900_CR47","doi-asserted-by":"crossref","unstructured":"Shao, C., Wu, X., & Feng, Y. (2022). One reference is not enough: Diverse distillation with reference selection for non-autoregressive translation. In: Carpuat, M., Marneffe, M., Ru\u00edz, I.V.M. (eds.) Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2022.naacl-main.277"},{"key":"2900_CR48","unstructured":"Shen, T., Ott, M., Auli, M., & Ranzato, M. (2019). Mixture models for diverse machine translation: Tricks of the trade. In K. Chaudhuri & R. Salakhutdinov (Eds.), Proceedings of the 36th International Conference on Machine Learning, ICML. Proceedings of Machine Learning Research."},{"key":"2900_CR49","doi-asserted-by":"crossref","unstructured":"Shen, X., Yuan, S., Sheng, H., Du, H., & Yu, X. (2023). Auslan-daily: Australian sign language translation for daily communication and news. In: Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track.","DOI":"10.52202\/075280-3527"},{"key":"2900_CR50","doi-asserted-by":"publisher","unstructured":"Shen, L., Zhan, H., Shen, X., Song, Y., & Zhao, X. (2021). Text is NOT enough: Integrating visual impressions into open-domain dialogue generation. In: Shen, H.T., Zhuang, Y., Smith, J.R., Yang, Y., C\u00e9sar, P., Metze, F., Prabhakaran, B. (eds.) MM \u201921: ACM Multimedia Conference, Virtual Event, China, October 20 - 24, 2021, pp. 4287\u20134296. ACM. https:\/\/doi.org\/10.1145\/3474085.3475568 .","DOI":"10.1145\/3474085.3475568"},{"key":"2900_CR51","doi-asserted-by":"crossref","unstructured":"Sheng, H., Shen, X., Du, H., Zhang, H., Huang, Z., & Yu, X. (2024). Ai empowered auslan learning for parents of deaf children and children of deaf adults. AI and Ethics, 1\u201311.","DOI":"10.1007\/s43681-024-00457-y"},{"key":"2900_CR52","doi-asserted-by":"crossref","unstructured":"Shi, B., Brentari, D., Shakhnarovich, G., & Livescu, K. (2022). Open-domain sign language translation learned from online video arXiv preprint arXiv:2205.12870.","DOI":"10.18653\/v1\/2022.emnlp-main.427"},{"key":"2900_CR53","doi-asserted-by":"publisher","unstructured":"Shu, R., Nakayama, H., & Cho, K. (2019). Generating diverse translations with sentence codes. In: Korhonen, A., Traum, D.R., M\u00e0rquez, L. (eds.) Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pp. 1823\u20131827. Association for Computational Linguistics. https:\/\/doi.org\/10.18653\/V1\/P19-1177 .","DOI":"10.18653\/V1\/P19-1177"},{"key":"2900_CR54","doi-asserted-by":"publisher","unstructured":"Su, J., Tan, Z., Xiong, D., Ji, R., Shi, X., & Liu, Y. (2017). Lattice-based recurrent neural network encoders for neural machine translation. In: Singh, S., Markovitch, S. (eds.) Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, February 4-9, 2017, San Francisco, California, USA, pp. 3302\u20133308. AAAI Press. https:\/\/doi.org\/10.1609\/AAAI.V31I1.10968 .","DOI":"10.1609\/AAAI.V31I1.10968"},{"key":"2900_CR55","doi-asserted-by":"publisher","unstructured":"Sun, Z., Huang, S., Wei, H., Dai, X., & Chen, J. (2020). Generating diverse translation by manipulating multi-head attention. In: The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, pp. 8976\u20138983. AAAI Press. https:\/\/doi.org\/10.1609\/AAAI.V34I05.6429.","DOI":"10.1609\/AAAI.V34I05.6429"},{"key":"2900_CR56","doi-asserted-by":"publisher","unstructured":"Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., and others. (2023). Llama 2: Open foundation and fine-tuned chat models. CoRR arXiv:2307.09288, https:\/\/doi.org\/10.48550\/ARXIV.2307.09288.","DOI":"10.48550\/ARXIV.2307.09288"},{"key":"2900_CR57","unstructured":"Vijayakumar, A. K., Cogswell, M., Selvaraju, R. R., Sun, Q., Lee, S., Crandall, D., & Batra, D. (2016). Diverse beam search: Decoding diverse solutions from neural sequence models arXiv preprint arXiv:1610.02424."},{"key":"2900_CR58","doi-asserted-by":"publisher","unstructured":"Vijayakumar, A.K., Cogswell, M., Selvaraju, R.R., Sun, Q., Lee, S., Crandall, D. J., & Batra, D. (2018). Diverse beam search for improved description of complex scenes. In: McIlraith, S. A., Weinberger, K. Q. (eds.) Proceedings of the Thirty-Second AAAI, pp. 7371\u20137379. AAAI Press. https:\/\/doi.org\/10.1609\/AAAI.V32I1.12340 .","DOI":"10.1609\/AAAI.V32I1.12340"},{"key":"2900_CR59","doi-asserted-by":"publisher","DOI":"10.1016\/J.NEUCOM.2023.126523","volume":"552","author":"Y Wei","year":"2023","unstructured":"Wei, Y., Yuan, S., Chen, M., Shen, X., Wang, L., Shen, L., & Yan, Z. (2023). Mpp-net: Multi-perspective perception network for dense video captioning. Neurocomputing, 552, Article 126523. https:\/\/doi.org\/10.1016\/J.NEUCOM.2023.126523","journal-title":"Neurocomputing"},{"key":"2900_CR60","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1007\/BF00992696","volume":"8","author":"RJ Williams","year":"1992","unstructured":"Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Mach. Learn., 8, 229\u2013256. https:\/\/doi.org\/10.1007\/BF00992696","journal-title":"Mach. Learn."},{"key":"2900_CR61","unstructured":"Wong, R., Camgoz, N. C., & Bowden, R. (2024). Sign2gpt: Leveraging large language models for gloss-free sign language translation. The Twelfth International Conference on Learning Representations."},{"key":"2900_CR62","doi-asserted-by":"crossref","unstructured":"Wu, X., Feng, Y., & Shao, C. (2020). Generating diverse translation from model distribution with dropout. In: Webber, B., Cohn, T., He, Y., Liu, Y. (eds.) Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020. Association for Computational Linguistics.","DOI":"10.18653\/v1\/2020.emnlp-main.82"},{"key":"2900_CR63","unstructured":"Wu, Y., Schuster, M., Chen, Z., Le, Q.V., Norouzi, M., Macherey, W., Krikun, M., and others. (2016). Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. CoRR arXiv:1609.08144."},{"key":"2900_CR64","doi-asserted-by":"publisher","unstructured":"Wu, L., Zhao, L., Qin, T., Lai, J., & Liu, T. (2017). Sequence prediction with unlabeled data by reward function learning. In: Sierra, C. (ed.) Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI 2017, Melbourne, Australia, August 19-25, 2017, pp. 3098\u20133104. ijcai.org. https:\/\/doi.org\/10.24963\/IJCAI.2017\/432.","DOI":"10.24963\/IJCAI.2017\/432"},{"key":"2900_CR65","doi-asserted-by":"crossref","unstructured":"Yang, A., Nagrani, A., Seo, P. H., Miech, A., Pont-Tuset, J., Laptev, I., Sivic, J., & Schmid, C. (2023). Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 10714\u201310726).","DOI":"10.1109\/CVPR52729.2023.01032"},{"key":"2900_CR66","doi-asserted-by":"crossref","unstructured":"Yin, K., & Read, J. (2020). Better sign language translation with stmc-transformer. Proceedings of the 28th International Conference on Computational Linguistics (pp. 5975\u20135989).","DOI":"10.18653\/v1\/2020.coling-main.525"},{"key":"2900_CR67","doi-asserted-by":"crossref","unstructured":"Yin, A., Zhong, T., Tang, L., Jin, W., Jin, T., & Zhao, Z. (2023). Gloss attention for gloss-free sign language translation. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 2551\u20132562.","DOI":"10.1109\/CVPR52729.2023.00251"},{"key":"2900_CR68","unstructured":"Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., & Xia, X. (2022). Glm-130b: An open bilingual pre-trained model arXiv preprint, arXiv:2210.02414."},{"key":"2900_CR69","doi-asserted-by":"publisher","unstructured":"Zhang, B., M\u00fcller, M., & Sennrich, R. (2023). SLTUNET: A simple unified model for sign language translation. CoRR. https:\/\/doi.org\/10.48550\/arXiv.2305.01778, arXiv:2305.01778.","DOI":"10.48550\/arXiv.2305.01778"},{"key":"2900_CR70","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Wu, W., Sun, W., Tu, D., Lu, W., Min, X., Chen, Y., & Zhai, G. (2023). Md-vqa: Multi-dimensional quality assessment for ugc live videos. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 1746\u20131755).","DOI":"10.1109\/CVPR52729.2023.00174"},{"key":"2900_CR71","doi-asserted-by":"publisher","first-page":"2662","DOI":"10.1109\/TMM.2021.3087006","volume":"24","author":"J Zhao","year":"2021","unstructured":"Zhao, J., Qi, W., Zhou, W., Duan, N., Zhou, M., & Li, H. (2021). Conditional sentence generation and cross-modal reranking for sign language translation. IEEE Transactions on Multimedia, 24, 2662\u20132672.","journal-title":"IEEE Transactions on Multimedia"},{"key":"2900_CR72","doi-asserted-by":"publisher","unstructured":"Zheng, R., Ma, M., & Huang, L. (2018). Multi-reference training with pseudo-references for neural translation and text generation. In: Riloff, E., Chiang, D., Hockenmaier, J., Tsujii, J. (eds.) Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pp. 3188\u20133197. Association for Computational Linguistics. https:\/\/doi.org\/10.18653\/V1\/D18-1357 .","DOI":"10.18653\/V1\/D18-1357"},{"key":"2900_CR73","doi-asserted-by":"crossref","unstructured":"Zheng, J., Wang, Y., Tan, C., Li, S., Wang, G., Xia, J., Chen, Y., & Li, S. Z. (2023). Cvt-slr: Contrastive visual-textual transformation for sign language recognition with variational alignment. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 23141\u201323150.","DOI":"10.1109\/CVPR52729.2023.02216"},{"key":"2900_CR74","doi-asserted-by":"publisher","first-page":"462","DOI":"10.1016\/j.neucom.2021.08.079","volume":"464","author":"J Zheng","year":"2021","unstructured":"Zheng, J., Chen, Y., Wu, C., Shi, X., & Kamal, S. M. (2021). Enhancing neural sign language translation by highlighting the facial expression information. Neurocomputing, 464, 462\u2013472. https:\/\/doi.org\/10.1016\/j.neucom.2021.08.079","journal-title":"Neurocomputing"},{"key":"2900_CR75","doi-asserted-by":"crossref","unstructured":"Zhou, B., Chen, Z., Clap\u00e9s, A., Wan, J., Liang, Y., Escalera, S., Lei, Z., & Zhang, D. (2023). Gloss-free sign language translation: Improving from visual-language pretraining. Proceedings of the IEEE\/CVF International Conference on Computer Vision (pp. 20871\u201320881).","DOI":"10.1109\/ICCV51070.2023.01908"},{"key":"2900_CR76","doi-asserted-by":"publisher","unstructured":"Zhou, H., Zhou, W., Qi, W., Pu, J., & Li, H. (2021). Improving sign language translation with monolingual data by sign back-translation. In: IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, Virtual, June 19-25, pp. 1316\u20131325. Computer Vision Foundation \/ IEEE, (2021). https:\/\/doi.org\/10.1109\/CVPR46437.2021.00137 . https:\/\/openaccess.thecvf.com\/content\/CVPR2021\/html\/Zhou_Improving_Sign_Language_Translation_With_Monolingual_Data_by_Sign_Back-Translation_CVPR_2021_paper.html","DOI":"10.1109\/CVPR46437.2021.00137"},{"key":"2900_CR77","doi-asserted-by":"crossref","unstructured":"Zhou, H., Zhou, W., Zhou, Y., & Li, H. (2021). Spatial-temporal multi-cue network for sign language recognition and translation. IEEE Transactions on Multimedia,24, 768\u2013779.","DOI":"10.1109\/TMM.2021.3059098"},{"key":"2900_CR78","doi-asserted-by":"publisher","unstructured":"Zuo, R., & Mak, B. (2022). $${\\rm C}^{2}$$slr: Consistency-enhanced continuous sign language recognition. In: IEEE\/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pp. 5121\u20135130. IEEE. https:\/\/doi.org\/10.1109\/CVPR52688.2022.00507 .","DOI":"10.1109\/CVPR52688.2022.00507"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02900-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-026-02900-5","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02900-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:41:53Z","timestamp":1782402113000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-026-02900-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":78,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["2900"],"URL":"https:\/\/doi.org\/10.1007\/s11263-026-02900-5","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"16 November 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 May 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 June 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"This study used publicly available sign language datasets, including PHOENIX14T, CSL-Daily, and How2Sign. In addition, human evaluation was conducted with native speakers and sign language experts for linguistic assessment only. No medical intervention, clinical procedure, or collection of sensitive personal data was involved. Ethical approval was not required for this study in accordance with the institutional guidelines of the authors\u2019 organizations.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical statement"}},{"value":"All authors have reviewed and approved the revised manuscript for publication.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"Informed consent was obtained from all individual participants involved in the human evaluation. For the publicly available datasets used in this study, participant consent was handled by the original dataset providers.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}}],"article-number":"329"}}