{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,15]],"date-time":"2026-01-15T17:47:41Z","timestamp":1768499261362,"version":"3.49.0"},"reference-count":33,"publisher":"World Scientific Pub Co Pte Ltd","issue":"07n08","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Artif. Intell. Tools"],"published-print":{"date-parts":[[2025,12]]},"abstract":"<jats:p>The rapid development of large language models (LLMs) has significantly benefited various industries. However, the psychological challenges experienced by left-behind children (LBC) due to a lack of companionship have received insufficient attention. To address this issue, we propose a multi-tone emotional voice synthesis framework that integrates speech recognition, LLMs and voice synthesis technologies. Specifically, an emotional psychology specialized knowledge and a psychological dialogue dataset tailored for LBC interactions. These resources enhance LLMs\u2019 capability to understand psychological issues and generate contextually appropriate responses that foster positive psychological development. Unlike conventional single-voice systems, our framework incorporates multi-tone audio samples through advanced speech recognition and synthesis techniques, enabling both diverse vocal outputs and efficient voice reconstruction. This approach significantly improves the naturalness and emotional expressiveness of synthesized speech. Experimental comparisons with ChatGLM3-6B and ChatGPT-3 demonstrate that our framework achieves superior performance in psychological question-answering tasks and voice expression. The results suggest its potential in providing effective psychological companionship for LBC, while offering practical implications for developing mental health support systems targeting this vulnerable population.<\/jats:p>","DOI":"10.1142\/s0218213025500228","type":"journal-article","created":{"date-parts":[[2025,12,11]],"date-time":"2025-12-11T16:51:24Z","timestamp":1765471884000},"source":"Crossref","is-referenced-by-count":0,"title":["LLM-Driven Psychological Companionship for LBC: A Multi-Tone Emotional Voice Synthesis Framework"],"prefix":"10.1142","volume":"34","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-5796-4937","authenticated-orcid":false,"given":"Shouyuan","family":"Qin","sequence":"first","affiliation":[{"name":"School of Computer and Information Engineering, Hubei Normal University, Huangshi, Hubei 435000, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-4696-5214","authenticated-orcid":false,"given":"Guixia","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer and Information Engineering, Hubei Normal University, Huangshi, Hubei 435000, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9347-3905","authenticated-orcid":false,"given":"Xuhui","family":"Xiong","sequence":"additional","affiliation":[{"name":"School of Computer and Information Engineering, Hubei Normal University, Huangshi, Hubei 435000, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2025,12,26]]},"reference":[{"issue":"5","key":"S0218213025500228BIB001","first-page":"786","volume":"37","author":"Zhao X.","year":"2023","journal-title":"China Science Foundation"},{"issue":"4","key":"S0218213025500228BIB002","first-page":"98","volume":"40","author":"Wang H.","year":"2023","journal-title":"Statistical Research"},{"key":"S0218213025500228BIB003","first-page":"59","author":"Fan X.","year":"2023","journal-title":"Journal of Beijing Normal University (Social Science Edition) (4)"},{"key":"S0218213025500228BIB004","first-page":"3","author":"Dai B.","year":"2022","journal-title":"China Special Education (3)"},{"key":"S0218213025500228BIB005","doi-asserted-by":"crossref","unstructured":"J. Maynez\n                      et al.\n                      , On faithfulness and factuality in abstractive summarization, arXiv:2005.00661 (2020).","DOI":"10.18653\/v1\/2020.acl-main.173"},{"key":"S0218213025500228BIB006","doi-asserted-by":"publisher","DOI":"10.1145\/3027063.3053246"},{"issue":"1","key":"S0218213025500228BIB007","first-page":"186","volume":"37","author":"Zhang A.","year":"2016","journal-title":"Small Microcomputer Systems"},{"key":"S0218213025500228BIB008","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2024.106762"},{"key":"S0218213025500228BIB009","doi-asserted-by":"crossref","unstructured":"Y. Wang\n                      et al.\n                      , Tacotron: Towards end-to-end speech synthesis, arXiv:1703.10135 (2017).","DOI":"10.21437\/Interspeech.2017-1452"},{"key":"S0218213025500228BIB010","first-page":"3171","volume":"32","author":"Ren Y.","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"S0218213025500228BIB011","unstructured":"A. Oord\n                      et al.\n                      , Wavenet: A generative model for raw audio, arXiv:1609.03499 (2016)."},{"key":"S0218213025500228BIB012","first-page":"17022","volume":"33","author":"Kong J.","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"S0218213025500228BIB013","first-page":"1877","volume":"33","author":"Brown T.","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"240","key":"S0218213025500228BIB014","first-page":"1","volume":"24","author":"Chowdhery A.","year":"2023","journal-title":"Journal of Machine Learning Research"},{"key":"S0218213025500228BIB015","unstructured":"A. Dubey\n                      et al.\n                      , The llama 3 herd of models, arXiv:2407.21783 (2024)."},{"key":"S0218213025500228BIB016","unstructured":"A. Yang\n                      et al.\n                      , Qwen2.5 technical report, arXiv:2412.15115 (2024)."},{"key":"S0218213025500228BIB017","doi-asserted-by":"publisher","DOI":"10.1111\/1911-3846.12832"},{"key":"S0218213025500228BIB018","first-page":"8748","volume-title":"Proc. Int. Conf. Machine Learning","author":"Radford A.","year":"2021"},{"key":"S0218213025500228BIB019","unstructured":"S. Wu\n                      et al.\n                      , Bloomberggpt: A large language model for finance, arXiv:2303.17564 (2023)."},{"key":"S0218213025500228BIB020","first-page":"5530","volume-title":"Proc. Int. Conf. Machine Learning","author":"Kim J.","year":"2021"},{"key":"S0218213025500228BIB021","unstructured":"J. Devlin\n                      et al.\n                      , Bert: Pre-training of deep bidirectional transformers for language understanding, arXiv:1810.04805 (2018)."},{"key":"S0218213025500228BIB022","doi-asserted-by":"crossref","unstructured":"H. Zhang\n                      et al.\n                      , Paddlespeech: An easy-to-use all-in-one speech toolkit, arXiv:2205.12007 (2022).","DOI":"10.18653\/v1\/2022.naacl-demo.12"},{"key":"S0218213025500228BIB023","first-page":"173","volume-title":"Proc. Int. Conf. Machine Learning","author":"Amodei D.","year":"2016"},{"key":"S0218213025500228BIB024","unstructured":"E. J. Hu\n                      et al.\n                      , Lora: Low-rank adaptation of large language models, arXiv:2106.09685 (2021)."},{"key":"S0218213025500228BIB025","first-page":"2790","volume-title":"Proc. Int. Conf. Machine Learning","author":"Houlsby N.","year":"2019"},{"key":"S0218213025500228BIB026","unstructured":"X. L. Li and P. Liang, Prefix-tuning: Optimizing continuous prompts for generation, arXiv:2101.00190 (2021)."},{"key":"S0218213025500228BIB027","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-short.8"},{"key":"S0218213025500228BIB028","doi-asserted-by":"crossref","unstructured":"X. Liu\n                      et al.\n                      , P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks, arXiv:2110.07602 (2021).","DOI":"10.18653\/v1\/2022.acl-short.8"},{"key":"S0218213025500228BIB029","doi-asserted-by":"publisher","DOI":"10.1145\/3422622"},{"key":"S0218213025500228BIB030","unstructured":"A. Zeng\n                      et al.\n                      , Glm-130b: An open bilingual pre-trained model, arXiv:2210.02414 (2022)."},{"key":"S0218213025500228BIB031","unstructured":"Z. Du\n                      et al.\n                      , Glm: General language model pretraining with autoregressive blank infilling, arXiv:2103.10360 (2021)."},{"key":"S0218213025500228BIB032","doi-asserted-by":"publisher","DOI":"10.1016\/j.dsp.2021.103058"},{"key":"S0218213025500228BIB033","doi-asserted-by":"publisher","DOI":"10.1186\/s12859-024-05729-2"}],"container-title":["International Journal on Artificial Intelligence Tools"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0218213025500228","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,15]],"date-time":"2026-01-15T02:51:35Z","timestamp":1768445495000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/10.1142\/S0218213025500228"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12]]},"references-count":33,"journal-issue":{"issue":"07n08","published-print":{"date-parts":[[2025,12]]}},"alternative-id":["10.1142\/S0218213025500228"],"URL":"https:\/\/doi.org\/10.1142\/s0218213025500228","relation":{},"ISSN":["0218-2130","1793-6349"],"issn-type":[{"value":"0218-2130","type":"print"},{"value":"1793-6349","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12]]},"article-number":"2550022"}}