{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T11:06:22Z","timestamp":1784027182664,"version":"3.55.0"},"reference-count":64,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T00:00:00Z","timestamp":1783987200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T00:00:00Z","timestamp":1783987200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Speech Technol"],"published-print":{"date-parts":[[2026,9]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Despite of swift advancements in the field of Artificial Intelligence, an effective immediate assistive solution for visually impaired individuals remains limited. Whereas, Deep Learning and Natural Language Processing (NLP) have attained significant gain in visual sympathetic and language generation but their addition into accessible systems for ecological perception is still underexplored. This dictates the development of intelligent frameworks capable of accurately rendering visual scenes and delivering meaningful descriptions to enhance liberation and situational awareness for visually impaired users. Consequently, we proposed a framework of three levels using the deep learning and NLP approaches to address aforementioned. It also developed a novel approach to learning the better relational features among the objects, scenes, and persons captured in an image and generated an accurate caption. Firstly, object detection algorithms are used to detect the objects in an image. The next level generated the image captions by maximizing the likelihood of the expected captions using deep learning. The third level, the text in the caption, is converted into a voice using NLP algorithms. The proposed model has substantially compared the existing state-of-the-art image captioning and voice conversion methods by experimenting with benchmark datasets such as Flickr 8k, Flickr 30k, COCO Caption, and RyanSpeech. The proposed model outperformed in terms of recall on Flickr 8K dataset @100 images has scored 67.25% and 69.10% rated on the same dataset @200 images. In addition, BLEU \u2212 1, 2 3, and 4 are scored 71.25%, 69.2%, 54.7%, and 40.8%.<\/jats:p>","DOI":"10.1007\/s10772-026-10281-w","type":"journal-article","created":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T10:39:54Z","timestamp":1784025594000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["An assistive framework for visually impaired using deep learning and natural language processing"],"prefix":"10.1007","volume":"29","author":[{"given":"V. Uma","family":"Maheswari","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rajanikanth","family":"Aluvalu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rajesh Kumar","family":"Dhanaraj","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mahmoud Ahmad","family":"Al-Khasawneh","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hariprasath","family":"Manoharan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shitharth","family":"Selvarajan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,7,14]]},"reference":[{"key":"10281_CR1","unstructured":"Alwakid, G. N., Humayun, M., & Ahmad, Z. (2025). Transforming disability into ability: An explainable vision-to-voice image captioning framework using Transformer models and edge computing. IEEE Access."},{"key":"10281_CR2","unstructured":"Babu, J. R., Sekharaiah, C., Kumar, G. M. (2015). A navigation tool for visually impaired persons. In 2015 (March) 2nd international conference on computing for sustainable global development (INDIACom) (pp. 938\u2013940). IEEE."},{"issue":"2","key":"10281_CR3","doi-asserted-by":"publisher","first-page":"335","DOI":"10.1007\/s00521-015-1846-7","volume":"27","author":"L. Bai","year":"2016","unstructured":"Bai, L., Li, K., Pei, J., & Jiang, S. (2016). Main objects interaction activity recognition in real images. Neural Computing & Applications, 27(2), 335\u2013348. https:\/\/doi.org\/10.1007\/s00521-015-1846-7","journal-title":"Neural Computing & Applications"},{"key":"10281_CR4","unstructured":"Banerjee, S., & Lavie, A. (2005). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL workshop on intrinsic and extrinsic evaluation measures for machine translation and\/or summarization (pp. 65\u201372)."},{"issue":"7","key":"10281_CR5","doi-asserted-by":"publisher","first-page":"2631","DOI":"10.1109\/TCYB.2018.2831447","volume":"49","author":"Y. Bin","year":"2019","unstructured":"Bin, Y., Yang, Y., Shen, F., Xie, N., Shen, H. T., & Li, X. (2019, Jul). Describing video with attention-based bidirectional L.S.T.M. IEEE Transactions on Cybernetics, 49(7), 2631\u20132641.","journal-title":"IEEE Trans. Cybern"},{"key":"10281_CR6","doi-asserted-by":"publisher","first-page":"845","DOI":"10.1109\/TMM.2021.3132724","volume":"25","author":"X. Cai","year":"2023","unstructured":"Cai, X., Liu, S., Han, J., Yang, L., Liu, Z., & Liu, T. (2023). Chestxraybert: A pretrained language model for chest radiology report summarization. IEEE Transactions on Multimedia, 25, 845\u2013855. https:\/\/doi.org\/10.1109\/TMM.2021.3132724","journal-title":"IEEE Transactions on Multimedia"},{"key":"10281_CR8","doi-asserted-by":"publisher","unstructured":"Castilho, S., Doherty, S., Gaspari, F., & Moorkens, J. (2018). Approaches to human and machine translation quality assessment. In J. Moorkens, S. Castilho, F. Gaspari, & S. Doherty (Eds.), Translation quality assessment (pp. 9\u201338). Springer. https:\/\/doi.org\/10.1007\/978-3-319-91241-7-2","DOI":"10.1007\/978-3-319-91241-7-2"},{"key":"10281_CR10","doi-asserted-by":"crossref","unstructured":"Chaccour, K., & Badr, G. (2016 ). Computer vision guidance system for indoor navigation of visually impaired people. In 2016 (September) IEEE 8th international conference on intelligent systems (IS) (pp. 449\u2013454). IEEE.","DOI":"10.1109\/IS.2016.7737460"},{"key":"10281_CR11","doi-asserted-by":"crossref","unstructured":"Chen, M., Hou, W., Ma, J., Wang, S., & Xiao, J. (2020). Non-parallel voice conversion with fewer labeled data by conditional generative adversarial networks. In Proceedings of Interspeech (pp. 4716\u20134720).","DOI":"10.21437\/Interspeech.2020-2162"},{"key":"10281_CR70","doi-asserted-by":"publisher","unstructured":"Chowdhury, I., Moeid, A., Hoque, E., Kabir, M. A., Hossain, M. S., & Islam, M. M. (2020). Designing and evaluating multimodal interactions for facilitating visual analysis with dashboards. IEEE Access, 9, 60\u201371. https:\/\/doi.org\/10.1109\/ACCESS.2020.3046623","DOI":"10.1109\/ACCESS.2020.3046623"},{"key":"10281_CR13","doi-asserted-by":"crossref","unstructured":"Dai, B., Fidler, S., Urtasun, R., & Lin, D. (2017). Towards diverse and natural image descriptions via a conditional GAN. In Proceedings of the IEEE international conference on computer vision (pp. 2970\u20132979) .","DOI":"10.1109\/ICCV.2017.323"},{"key":"10281_CR14","doi-asserted-by":"publisher","first-page":"116942","DOI":"10.1109\/ACCESS.2022.3219455","volume":"10","author":"P. Danenas","year":"2022","unstructured":"Danenas, P., & Skersys, T. (2022). Exploring natural language processing in model-to-model transformations. IEEE Access, 10, 116942\u2013116958. https:\/\/doi.org\/10.1109\/ACCESS.2022.3219455","journal-title":"IEEE Access"},{"key":"10281_CR15","doi-asserted-by":"crossref","unstructured":"Doddington, G. (2002). Automatic evaluation of machine translation quality using N-gram co-occurrence statistics. In HLT\u201902: Proceedings of the second international conference on human language technology (pp. 138\u2013144).","DOI":"10.3115\/1289189.1289273"},{"key":"10281_CR16","doi-asserted-by":"crossref","unstructured":"Fattahi, J., & Mejri, M. (2021). SpaML: A bimodal ensemble learning spam detector based on NLP techniques. In 2021 (January) IEEE 5th international conference on cryptography, security and privacy (CSP) (pp. 107\u2013112). IEEE.","DOI":"10.1109\/CSP51677.2021.9357595"},{"key":"10281_CR17","doi-asserted-by":"publisher","first-page":"3709","DOI":"10.1109\/TIFS.2020.2997134","volume":"15","author":"Q. Feng","year":"2020","unstructured":"Feng, Q., He, D., Liu, Z., Wang, H., & Choo, K. K. R. (2020). SecureNLP: A system for multi-party privacy-preserving natural language processing. IEEE Transactions on Information Forensics and Security, 15, 3709\u20133721. https:\/\/doi.org\/10.1109\/TIFS.2020.2997134","journal-title":"IEEE Transactions on Information Forensics and Security"},{"key":"10281_CR18","doi-asserted-by":"crossref","unstructured":"Firdus, S., Ahmad, W. F. W., & Janier, J. B. (2012, June). Development of audio video describer using narration to visualize movie film for blind and visually impaired children. In 2012 international conference on computer & information science (ICCIS) (Vol. 2, pp. 1068\u20131072). IEEE.","DOI":"10.1109\/ICCISci.2012.6297184"},{"issue":"12","key":"10281_CR19","doi-asserted-by":"publisher","first-page":"2321","DOI":"10.1109\/TPAMI.2016.2642953","volume":"39","author":"K. Fu","year":"2017","unstructured":"Fu, K., Jin, J., Cui, R., Sha, F., & Zhang, C. (2017, Dec). Aligning where to see and what to tell: Image captioning with region-based attention and scene-specific contexts. IEEE Transactions on Pattern Analysis & Machine Intelligence, 39(12), 2321\u20132334. https:\/\/doi.org\/10.1109\/TPAMI.2016.2642953","journal-title":"IEEE Transactions on Pattern Analysis & Machine Intelligence"},{"key":"10281_CR20","doi-asserted-by":"crossref","unstructured":"Gagnon, L., Chapdelaine, C., Byrns, D., Foucher, S., Heritier, M., & Gupta, V. (2010). A computer-vision-assisted system for videodescription scripting. In 2010 (June) IEEE computer society conference on computer vision and pattern recognition-workshops (pp. 41\u201348).","DOI":"10.1109\/CVPRW.2010.5543575"},{"key":"10281_CR21","doi-asserted-by":"crossref","unstructured":"Gupta, P., & Katal, N. (2023). Deep learning based automatic image caption generation for visually impaired people. In Intelligent systems and applications in computer vision (pp. 141\u2013157). CRC Press.","DOI":"10.1201\/9781003453406-12"},{"issue":"8","key":"10281_CR22","doi-asserted-by":"publisher","first-page":"7441","DOI":"10.1109\/TCYB.2020.3041595","volume":"52","author":"L. Huo","year":"2022","unstructured":"Huo, L., Bai, L., & Zhou, S. M. (2022). Automatically generating natural language descriptions of images by a deep hierarchical framework. IEEE Transactions on Cybernetics, 52(8), 7441\u20137452. https:\/\/doi.org\/10.1109\/TCYB.2020.3041595","journal-title":"IEEE Transactions on Cybernetics"},{"key":"10281_CR23","doi-asserted-by":"crossref","unstructured":"Janier, J. B., Ahmad, W. F. W., & Firdus, S. B. (2013, January). Representing visual content of movie cartoons through narration for the visually impaired. In 2013 international conference on computer applications technology (ICCAT) (pp. 1\u20136). IEEE.","DOI":"10.1109\/ICCAT.2013.6522042"},{"key":"10281_CR24","unstructured":"Jayanthi, K., Raghavan, H., Murarka, S., & Murthy, H. A. (2023). Design and Development of a text-to-speech synthesizer for Indian languages."},{"key":"10281_CR25","doi-asserted-by":"crossref","unstructured":"Kain, A., & Macon, M. W. (1998 ). Spectral voice conversion for text-to-speech synthesis. In Proceedings of ICASSP (May) (pp. 285\u2013288, Seattle, WA.","DOI":"10.1109\/ICASSP.1998.674423"},{"key":"10281_CR26","doi-asserted-by":"crossref","unstructured":"Kotani, G., Saito, D., & Minematsu, N. (2017, December). Voice conversion based on deep neural networks for time-variant linear transformations. In 2017 Asia-Pacific signal and I information processing association annual summit and conference (APSIPA ASC) (pp. 1259\u20131262). IEEE.","DOI":"10.1109\/APSIPA.2017.8282216"},{"key":"10281_CR27","doi-asserted-by":"publisher","unstructured":"Kulkarni, G., , V., Ordonez, V., Dhar, S., Li, S., Choi, Y., & Berg, A. C. (2013, Dec). BabyTalk: Understanding and generating simple image descriptions. IEEE Transactions on Pattern Analysis & Machine Intelligence, 35(12), 2891\u20132903. https:\/\/doi.org\/10.1109\/TPAMI.2012.162","DOI":"10.1109\/TPAMI.2012.162"},{"issue":"10","key":"10281_CR28","doi-asserted-by":"publisher","first-page":"351","DOI":"10.1162\/tacl_a_00188","volume":"2","author":"P. Kuznetsova","year":"2014","unstructured":"Kuznetsova, P., Ordonez, V., Berg, T. L., & Choi, Y. (2014). TreeTalk: Composition and compression of trees for image descriptions. Transactions of the Association for Computational Linguistics, 2(10), 351\u2013362. https:\/\/doi.org\/10.1162\/tacl_a_00188","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"10281_CR29","doi-asserted-by":"crossref","unstructured":"Lavie, A., & Agarwal, A. (2007). METEOR: An automatic metric for MT evaluation with high levels of correlation with human judgments. In Proceedings of the workshop on statistical machine translation (pp. 228\u2013231, June, Prague, Czech Republic.","DOI":"10.3115\/1626355.1626389"},{"issue":"2","key":"10281_CR30","doi-asserted-by":"publisher","first-page":"913","DOI":"10.1109\/TCYB.2019.2914351","volume":"51","author":"X. Li","year":"2021","unstructured":"Li, X., Yuan, A., & Lu, X. (2021). Vision-to-language tasks based on attributes and attention mechanism. IEEE Transactions on Cybernetics, 51(2), 913\u2013926. https:\/\/doi.org\/10.1109\/TCYB.2019.2914351","journal-title":"IEEE Transactions on Cybernetics"},{"key":"10281_CR31","unstructured":"Li, X., Yuan, A., & Lu, X. (2019, May 17). Vision-to-language tasks based on attributes and attention mechanism. IEEE Transactons on Cybernetics. Advance online publication."},{"key":"10281_CR32","doi-asserted-by":"crossref","unstructured":"Lin, C.-Y., & Hovy, E. (2003). Automatic evaluation of summaries using N-gram co-occurrence statistics. In Proceedings of the 2003 human language technology conference of the North American chapter of the association for computational linguistics (pp. 150\u2013157).","DOI":"10.3115\/1073445.1073465"},{"key":"10281_CR33","doi-asserted-by":"publisher","first-page":"2967","DOI":"10.1109\/TASLP.2020.3034994","volume":"28","author":"H. T. Luong","year":"2020","unstructured":"Luong, H. T., & Yamagishi, J. (2020). Nautilus: A versatile voice cloning system. IEEE\/ACM Transactions on Audio, Speech, and Language Processing, 28, 2967\u20132981. https:\/\/doi.org\/10.1109\/TASLP.2020.3034994","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"10281_CR34","unstructured":"Mao, J., Xu, W., Yang, Y., Wang, J., & Yuille, A. L. (2015). Deep captioning with multimodal recurrent neural networks (m-RNN). In Proceedings of the international conference on learning representations, May, San Diego, CA, USA."},{"issue":"5","key":"10281_CR35","doi-asserted-by":"publisher","first-page":"969","DOI":"10.1109\/JSTSP.2020.2994523","volume":"14","author":"Z. Mi","year":"2020","unstructured":"Mi, Z., Jiang, X., Sun, T., & Xu, K. (2020). GAN-generated image detection with self-attention mechanism against GAN generator defect. IEEE Journal of Selected Topics in Signal Processing, 14(5), 969\u2013981. https:\/\/doi.org\/10.1109\/JSTSP.2020.2994523","journal-title":"IEEE Journal of Selected Topics in Signal Processing"},{"key":"10281_CR36","unstructured":"Mitchell, M., et al. (2012). MIDGE: Generating image descriptions from computer vision detections. In Proceedings of the 13th Conference of the European chapter of the association for computational linguistics (pp. 747\u2013756)."},{"key":"10281_CR37","doi-asserted-by":"crossref","unstructured":"Mohanraj, P., Rajasekar, T., Sivaelango, N., Karthickraja, V. S., & Vignesh, N. (2024). Wearable device for visually impaired using deep learning. In 2024 (June) 3rd international conference on applied artificial intelligence and computing (ICAAIC) (pp. 348\u2013352). IEEE.","DOI":"10.1109\/ICAAIC60222.2024.10575667"},{"key":"10281_CR38","doi-asserted-by":"crossref","unstructured":"Nakamura, S., & Shikano, K. (1989, May). Speaker adaptation applied to HMM and neural networks. In Proceedings of the IEEE international conference on acoustics, speech, and signal processing (ICASSP) (pp. 89\u201392), May, Glasgow, U.K.","DOI":"10.1109\/ICASSP.1989.266370"},{"key":"10281_CR39","doi-asserted-by":"crossref","unstructured":"Nie\u00dfen, S., Och, F. J., Leusch, G., & Ney, H. (2000). An evaluation tool for machine translation: Fast evaluation for MT research. In Proceedings of the second international conference on language resources and evaluation (LREC).","DOI":"10.63317\/56t9t4rchhar"},{"key":"10281_CR40","doi-asserted-by":"publisher","unstructured":"Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting on association for computational linguistics. https:\/\/doi.org\/10.3115\/1073083.1073135","DOI":"10.3115\/1073083.1073135"},{"key":"10281_CR42","doi-asserted-by":"publisher","first-page":"2364","DOI":"10.1109\/TASLP.2020.3012060","volume":"28","author":"Y. Qin","year":"2020","unstructured":"Qin, Y., Qi, F., Ouyang, S., Liu, Z., Yang, C., Wang, Y., Liu, Q., & Sun, M. (2020). Improving sequence modeling ability of recurrent neural networks via sememes. IEEE\/ACM Transactions on Audio, Speech, and Language Processing, 28, 2364\u20132373. https:\/\/doi.org\/10.1109\/TASLP.2020.3012060","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"10281_CR43","doi-asserted-by":"crossref","unstructured":"Redmon, J., & Farhadi, A. (2017). YOLO9000: Better, faster, stronger. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7263\u20137271).","DOI":"10.1109\/CVPR.2017.690"},{"key":"10281_CR44","doi-asserted-by":"crossref","unstructured":"Rocha Fa\u00e7anha, A., Caetano de Oliveira, A., Vinicius de Andrade Lima, M., Viana, W., & S\u00e1nchez, J. (2016). Audio description of videos for people with visual disabilities. In Universal access in human-computer interaction. Users and context diversity: 10th international conference, U.A.H.C.I. 2016, Held as Part of HCI International 2016 (pp. 505\u2013515), 17\u201322 July 2016, Toronto, ON, Canada.","DOI":"10.1007\/978-3-319-40238-3_48"},{"key":"10281_CR45","doi-asserted-by":"publisher","first-page":"591","DOI":"10.1109\/TIP.2019.2930176","volume":"29","author":"X. Rong","year":"2020","unstructured":"Rong, X., Yi, C., & Tian, Y. (2020). Unambiguous scene text segmentation with referring expression comprehension. IEEE Transactions on Image Processing, 29, 591\u2013601. https:\/\/doi.org\/10.1109\/TIP.2019.2930176","journal-title":"IEEE Transactions on Image Processing"},{"issue":"20","key":"10281_CR46","doi-asserted-by":"publisher","first-page":"59413","DOI":"10.1007\/s11042-023-17849-7","volume":"83","author":"K. M. Safiya","year":"2023","unstructured":"Safiya, K. M., & Pandian, R. (2023). A real-time image captioning framework using computer vision to help the visually impaired. Multimedia Tools & Applications, 83(20), 59413\u201359438. https:\/\/doi.org\/10.1007\/s11042-023-17849-7","journal-title":"Multimedia Tools & Applications"},{"key":"10281_CR47","doi-asserted-by":"publisher","first-page":"52926","DOI":"10.1109\/ACCESS.2021.3069205","volume":"9","author":"K. C. Shahira","year":"2021","unstructured":"Shahira, K. C., & Lijiya, A. (2021). Towards assisting the visually impaired: A review on techniques for decoding the visual data from chart images. IEEE Access, 9, 52926\u201352943. https:\/\/doi.org\/10.1109\/ACCESS.2021.3069205","journal-title":"IEEE Access"},{"key":"10281_CR48","doi-asserted-by":"publisher","first-page":"3552","DOI":"10.1109\/TIP.2023.3287038","volume":"32","author":"P. Shivakumara","year":"2023","unstructured":"Shivakumara, P., Banerjee, A., Pal, U., Nandanwar, L., Lu, T., & Liu, C. L. (2023). A new language-independent deep CNN for scene text detection and style transfer in social media images. IEEE Transactions on Image Processing, 32, 3552\u20133566. https:\/\/doi.org\/10.1109\/TIP.2023.3287038","journal-title":"IEEE Transactions on Image Processing"},{"issue":"2","key":"10281_CR50","doi-asserted-by":"publisher","first-page":"131","DOI":"10.1109\/89.661472","volume":"6","author":"Y. Stylianou","year":"1998","unstructured":"Stylianou, Y., Capp\u00e9, O., & Moulines, E. (1998, Mar). Continuous probabilistic transform for voice conversion. IEEE Transactions on Speech and Audio Processing, 6(2), 131\u2013142. https:\/\/doi.org\/10.1109\/89.661472","journal-title":"IEEE Transactions on Speech and Audio Processing"},{"key":"10281_CR52","doi-asserted-by":"crossref","unstructured":"Sun, L., Kang, S., Li, K., & Meng, H. (2015). Voice conversion using deep bidirectional long short-term memory based recurrent neural networks. In Proceedings of the international conference on acoustics, speech, and signal processing (pp. 4869\u20134873).","DOI":"10.1109\/ICASSP.2015.7178896"},{"issue":"5","key":"10281_CR53","doi-asserted-by":"publisher","first-page":"932","DOI":"10.1109\/TASL.2010.2041688","volume":"18","author":"J. Tao","year":"2010","unstructured":"Tao, J., Zhang, M., Nurminen, J., Tian, J., & Wang, X. (2010). Supervisory data alignment for text-independent voice conversion. IEEE Transactions on Audio, Speech, and Language Processing, 18(5), 932\u2013943. https:\/\/doi.org\/10.1109\/TASL.2010.2041688","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"issue":"8","key":"10281_CR54","doi-asserted-by":"publisher","first-page":"2222","DOI":"10.1109\/TASL.2007.907344","volume":"15","author":"T. Toda","year":"2007","unstructured":"Toda, T., Black, A. W., & Tokuda, K. (2007, Nov). Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory. IEEE Transactions on Audio, Speech, and Language Processing, 15(8), 2222\u20132235. https:\/\/doi.org\/10.1109\/TASL.2007.907344","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"key":"10281_CR55","doi-asserted-by":"crossref","unstructured":"Toda, T., Saruwatari, H., & Shikao, K. (2001). Voice conversion algorithm based on Gaussian mixture model with dynamic frequency warping of straight spectrum. In Proceedings of the international conference on acoustics, speech, and signal processing (ICASSP)  (pp. 841\u2013944).","DOI":"10.1109\/ICASSP.2001.941046"},{"issue":"5","key":"10281_CR56","doi-asserted-by":"publisher","first-page":"816","DOI":"10.1093\/ietisy\/e90-d.5.816","volume":"E90-D","author":"T. Toda","year":"2007","unstructured":"Toda, T., & Tokuda, K. (2007). A speech parameter generation algorithm considering global variance for HMM-based speech synthesis. IEICE Transactions on Information and Systems, E90-D(5), 816\u2013824. https:\/\/doi.org\/10.1093\/ietisy\/e90-d.5.816","journal-title":"IEICE Transactions on Information and Systems"},{"key":"10281_CR57","unstructured":"Turian, J. P., Shen, L., & Melamed, I. D. (2003). Evaluation of machine translation and its evaluation. In Proceedings of MT Summit IX (pp. 386\u2013393, New Orleans, USA."},{"key":"10281_CR69","unstructured":"Vinyals, O., Blundell, C., Lillicrap, T., Kavukcuoglu, K., & Wierstra, D. (2016). Matching networks for one shot learning. Advances in Neural Information Processing Systems, 29. https:\/\/arxiv.org\/abs\/1606.04080"},{"issue":"4","key":"10281_CR58","doi-asserted-by":"publisher","first-page":"652","DOI":"10.1109\/TPAMI.2016.2587640","volume":"39","author":"O. Vinyals","year":"2017","unstructured":"Vinyals, O., Toshev, A., Bengio, S., & Erhan, D. (2017). Vinyals: Lessons learned from the 2015 M.S.C.O.C.O. image captioning challenge. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(4), 652\u2013663.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"10281_CR59","unstructured":"Wang, Q., & Chan, A. B. (2018). CNN+ CNN: Convolutional decoders for image captioning."},{"issue":"4","key":"10281_CR60","doi-asserted-by":"publisher","first-page":"1109","DOI":"10.1109\/TASL.2006.876112","volume":"14","author":"C. H. Wu","year":"2006","unstructured":"Wu, C. H., Hsia, C. C., Liu, T. H., & Wang, J. F. (2006). Voice conversion using duration-embedded bi-HMMs for expressive speech synthesis. IEEE Transactions on Audio, Speech, and Language Processing, 14(4), 1109\u20131116. https:\/\/doi.org\/10.1109\/TASL.2006.876112","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"key":"10281_CR61","doi-asserted-by":"crossref","unstructured":"Xie, J., Cai, Y., Chen, J., Xu, R., Wang, J., & Li, Q. (2024). Knowledge-augmented visual question answering with natural language explanation. IEEE Transactions on Image Processing.","DOI":"10.1109\/TIP.2024.3379900"},{"key":"10281_CR62","doi-asserted-by":"publisher","first-page":"9627","DOI":"10.1109\/TIP.2020.3028651","volume":"29","author":"M. Yang","year":"2020","unstructured":"Yang, M., Liu, J., Shen, Y., Zhao, Z., Chen, X., Wu, Q., & Li, C. (2020). An ensemble of generation- and retrieval-based image captioning with dual generator generative adversarial network. IEEE Transactions on Image Processing, 29, 9627\u20139640. https:\/\/doi.org\/10.1109\/TIP.2020.3028651","journal-title":"IEEE Transactions on Image Processing"},{"key":"10281_CR63","doi-asserted-by":"crossref","unstructured":"Yao, T., Pan, Y., Li, Y., & Mei, T. (2018). Exploring visual relationship for image captioning. In Proceedings of the European conference on computer vision (ECCV) (pp. 684\u2013699).","DOI":"10.1007\/978-3-030-01264-9_42"},{"key":"10281_CR64","doi-asserted-by":"publisher","first-page":"92","DOI":"10.1109\/TMM.2020.2976552","volume":"23","author":"J. Zhang","year":"2021","unstructured":"Zhang, J., Mei, K., Zheng, Y., & Fan, J. (2021). Integrating part of speech guidance for image captioning. IEEE Transactions on Multimedia, 23, 92\u2013104. https:\/\/doi.org\/10.1109\/TMM.2020.2976552","journal-title":"IEEE Transactions on Multimedia"},{"key":"10281_CR65","doi-asserted-by":"crossref","unstructured":"Zhang, M., Tao, J., Nurminen, J., Tian, J., & Wang, X. (2008). Phonetic anchor-based state mapping for text-independent voice conversion. In 2008 (October) 9th international conference on signal processing (pp. 723\u2013727). IEEE.","DOI":"10.1109\/ICOSP.2008.4697232"},{"key":"10281_CR66","doi-asserted-by":"publisher","first-page":"2662","DOI":"10.1109\/TMM.2021.3087006","volume":"24","author":"J. Zhao","year":"2022","unstructured":"Zhao, J., Qi, W., Zhou, W., Duan, N., Zhou, M., & Li, H. (2022). Conditional sentence generation and cross-modal reranking for sign language translation. IEEE Transactions on Multimedia, 24, 2662\u20132672. https:\/\/doi.org\/10.1109\/TMM.2021.3087006","journal-title":"IEEE Transactions on Multimedia"},{"key":"10281_CR67","unstructured":"Zhao, Y., Huang, W. C., Tian, X., Yamagishi, J., Das, R. K., Kinnunen, T., Ling, Z., & Toda, T. (2020). Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion. arXiv preprint arXiv: 2008.12527."},{"key":"10281_CR68","doi-asserted-by":"crossref","unstructured":"Zheng, R. C., Ai, Y., & Ling, Z. H. (2024). Incorporating ultrasound tongue images for audio-visual speech enhancement. IEEE\/ACM Transactions on Audio, Speech, and Language Processing.","DOI":"10.1109\/TASLP.2024.3361376"}],"container-title":["International Journal of Speech Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10772-026-10281-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10772-026-10281-w","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10772-026-10281-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T10:40:02Z","timestamp":1784025602000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10772-026-10281-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,14]]},"references-count":64,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,9]]}},"alternative-id":["10281"],"URL":"https:\/\/doi.org\/10.1007\/s10772-026-10281-w","relation":{},"ISSN":["1381-2416","1572-8110"],"issn-type":[{"value":"1381-2416","type":"print"},{"value":"1572-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,14]]},"assertion":[{"value":"27 January 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"11 June 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 July 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"58"}}