{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T15:15:40Z","timestamp":1787066140084,"version":"3.56.0"},"reference-count":47,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2025,8,27]],"date-time":"2025-08-27T00:00:00Z","timestamp":1756252800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>This paper presents a modular framework for character-coherent, emotion-aware role-playing dialogue with large language models (LLMs), centered on a novel Verifiable Emotion Reward (VER) objective. We introduce VER as a reinforcement-style signal derived from frozen emotion classifiers to provide both turn-level and dialogue-level alignment, effectively mitigating emotional drift across long interactions. To amplify VER\u2019s benefits, we construct Character-Coherent Dialogues (CHARCO), a large-scale multi-turn dataset of over 230,000 dialogues, richly annotated with persona profiles, semantic contexts, and ten emotion labels. Our experiments show that fine-tuning LLMs on CHARCO significantly enhances VER\u2019s impact, driving marked improvements in emotional consistency, role fidelity, and dialogue coherence. Through the evaluation that integrates lexical diversity metrics, automatic scoring with GPT-4, and human assessments, we demonstrate that the collaboration between a purpose-built multi-turn dataset and the VER objective leads to significant advancements in the field of persona-aligned conversational agents.<\/jats:p>","DOI":"10.3390\/info16090738","type":"journal-article","created":{"date-parts":[[2025,8,27]],"date-time":"2025-08-27T09:14:05Z","timestamp":1756286045000},"page":"738","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Enhancing Character-Coherent Role-Playing Dialogue with a Verifiable Emotion Reward"],"prefix":"10.3390","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-4298-5885","authenticated-orcid":false,"given":"Junqiao","family":"Wang","sequence":"first","affiliation":[{"name":"Sichuan University Pittsburgh Institute, Sichuan University, Chengdu 610207, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-5046-3771","authenticated-orcid":false,"given":"Kunyu","family":"Wu","sequence":"additional","affiliation":[{"name":"Sichuan University Pittsburgh Institute, Sichuan University, Chengdu 610207, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1611-1114","authenticated-orcid":false,"given":"Yuqi","family":"Ouyang","sequence":"additional","affiliation":[{"name":"College of Computer Science, Sichuan University, Chengdu 610207, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,8,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Raikov, A., Giretti, A., Pirani, M., Spalazzi, L., and Guo, M. (2024). Accelerating human\u2013computer interaction through convergent conditions for LLM explanation. Front. Artif. Intell., 7.","DOI":"10.3389\/frai.2024.1406773"},{"key":"ref_2","unstructured":"Sowmiya, R., Revathi, P., Ragunath, D., Gokila, P., and Kalaivani, T. (2024, January 3\u20135). Multi-Modal LLM Driven Computer Interface. Proceedings of the 2024 8th International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud)(I-SMAC), Kirtipur, Nepal."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"260","DOI":"10.1007\/s10462-024-10888-y","article-title":"Large language models (LLMs): Survey, technical frameworks, and future challenges","volume":"57","author":"Kumar","year":"2024","journal-title":"Artif. Intell. Rev."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1007\/s11280-024-01276-1","article-title":"When large language models meet personalization: Perspectives of challenges and opportunities","volume":"27","author":"Chen","year":"2024","journal-title":"World Wide Web"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Nazi, Z.A., and Peng, W. (2024). Large language models in healthcare and medical domain: A review. Informatics, 11.","DOI":"10.3390\/informatics11030057"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Bolpagni, M., and Gabrielli, S. (2025). Development of a comprehensive evaluation scale for LLM-powered counseling chatbots (CES-LCC) using the Edelphi method. Informatics, 12.","DOI":"10.20944\/preprints202501.1621.v1"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Pinto-Bernal, M., Biondina, M., and Belpaeme, T. (2025). Designing Social Robots with LLMs for Engaging Human Interaction. Appl. Sci., 15.","DOI":"10.3390\/app15116377"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Jedrzejczak, W.W., and Kobosko, J. (2025). Do Chatbots Exhibit Personality Traits? A Comparison of ChatGPT and Gemini Through Self-Assessment. Information, 16.","DOI":"10.3390\/info16070523"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Klinkert, L.J., Buongiorno, S., and Clark, C. (2024, January 18\u201322). Evaluating the efficacy of LLMs to emulate realistic human personalities. Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, Lexington, KY, USA.","DOI":"10.1609\/aiide.v20i1.31867"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Chen, Y.C., Lee, S.H., Sheu, H., Lin, S.H., Hu, C.C., Fu, S.C., Yang, C.P., and Lin, Y.C. (2025). Enhancing responses from large language models with role-playing prompts: A comparative study on answering frequently asked questions about total knee arthroplasty. BMC Med. Inform. Decis. Mak., 25.","DOI":"10.1186\/s12911-025-03024-5"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Feng, Q., Xie, Q., Wang, X., Li, Q., Zhang, Y., Feng, R., Zhang, T., and Gao, S. (May, January 29). EmoCharacter: Evaluating the Emotional Fidelity of Role-Playing Agents in Dialogues. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Albuquerque, NM, USA.","DOI":"10.18653\/v1\/2025.naacl-long.316"},{"key":"ref_12","unstructured":"Che, W., Nabende, J., Shutova, E., and Pilehvar, M.T. (August, January 27). Beyond Dialogue: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Dan, Y., Zhou, J., Chen, Q., Tian, J., and He, L. (August, January 27). P-React: Synthesizing Topic-Adaptive Reactions of Personality Traits via Mixture of Specialized LoRA Experts. Proceedings of the Findings of the Association for Computational Linguistics: ACL 2025, Vienna, Austria.","DOI":"10.18653\/v1\/2025.findings-acl.328"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Huang, L., Lan, H., Sun, Z., Shi, C., and Bai, T. (2024, January 11\u201312). Emotional RAG: Enhancing role-playing agents through emotional retrieval. Proceedings of the 2024 IEEE International Conference on Knowledge Graph (ICKG), Abu Dhabi, United Arab Emirates.","DOI":"10.1109\/ICKG63256.2024.00023"},{"key":"ref_15","unstructured":"Che, W., Nabende, J., Shutova, E., and Pilehvar, M.T. (August, January 27). Context-Aware Sentiment Forecasting via LLM-based Multi-Perspective Role-Playing Agents. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria."},{"key":"ref_16","unstructured":"Gomes, A., Brito, E., Morais, L., and Ferreira, N. (2025). How do Data Journalists Design Maps to Tell Stories?. arXiv, Available online: http:\/\/arxiv.org\/abs\/2508.10903."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yu, T., Shi, K., Zhao, Z., and Penn, G. (2025, January 4). Multi-Agent Based Character Simulation for Story Writing. Proceedings of the Fourth Workshop on Intelligent and Interactive Writing Assistants (In2Writing 2025), Albuquerque, NM, USA.","DOI":"10.18653\/v1\/2025.in2writing-1.9"},{"key":"ref_18","unstructured":"Zhang, P., An, S., Qiao, L., Yu, Y., Chen, J., Wang, J., Yin, D., Sun, X., and Zhang, K. (August, January 27). RolePlot: A Systematic Framework for Evaluating and Enhancing the Plot-Progression Capabilities of Role-Playing Agents. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Ding, N., Chen, Y., Xu, B., Qin, Y., Zheng, Z., Hu, S., Liu, Z., Sun, M., and Zhou, B. (2023). Enhancing chat language models by scaling high-quality instructional conversations. arXiv.","DOI":"10.18653\/v1\/2023.emnlp-main.183"},{"key":"ref_20","unstructured":"Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T.B. (2025, July 24). Stanford Alpaca: An Instruction-Following LLaMA Model, 2023. Available online: https:\/\/github.com\/tatsu-lab\/stanford_alpaca."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Xu, C., Guo, D., Duan, N., and McAuley, J. (2023). Baize: An open-source chat model with parameter-efficient tuning on self-chat data. arXiv.","DOI":"10.18653\/v1\/2023.emnlp-main.385"},{"key":"ref_22","unstructured":"Che, W., Nabende, J., Shutova, E., and Pilehvar, M.T. (August, January 27). KokoroChat: A Japanese Psychological Counseling Dialogue Dataset Collected via Role-Playing by Trained Counselors. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria."},{"key":"ref_23","unstructured":"Tao, M., Liang, X., Shi, T., Yu, L., and Xie, Y. (2024). RoleCraft-GLM: Advancing Personalized Role-Playing in Large Language Models. arXiv, Available online: http:\/\/arxiv.org\/abs\/2401.09432."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Kim, H., Hessel, J., Jiang, L., West, P., Lu, X., Yu, Y., Zhou, P., Bras, R.L., Alikhani, M., and Kim, G. (2022). SODA: Million-scale Dialogue Distillation with Social Commonsense Contextualization. arXiv.","DOI":"10.18653\/v1\/2023.emnlp-main.799"},{"key":"ref_25","unstructured":"Ji, Y. (2023). Exploring the Impact of Instruction Data Scaling on Large Language Models: An Empirical Study on Real-World Use Cases. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Tu, Q., Fan, S., Tian, Z., and Yan, R. (2024). Charactereval: A chinese benchmark for role-playing conversational agent evaluation. arXiv.","DOI":"10.18653\/v1\/2024.acl-long.638"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Wang, Z.M., Peng, Z., Que, H., Liu, J., Zhou, W., Wu, Y., Guo, H., Gan, R., Ni, Z., and Yang, J. (2023). Rolellm: Benchmarking, eliciting, and enhancing role-playing abilities of large language models. arXiv.","DOI":"10.18653\/v1\/2024.findings-acl.878"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002, January 6\u201312). BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, Philadelphia, PA, USA.","DOI":"10.3115\/1073083.1073135"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"71","DOI":"10.3115\/1073445.1073465","article-title":"Automatic evaluation of summaries using N-gram co-occurrence statistics","volume":"Volume 1","author":"Lin","year":"2003","journal-title":"Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"381","DOI":"10.3758\/BRM.42.2.381","article-title":"MTLD, vocd-D, and HD-D: A validation study of sophisticated approaches to lexical diversity assessment","volume":"42","author":"McCarthy","year":"2010","journal-title":"Behav. Res. Methods"},{"key":"ref_31","unstructured":"Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., and Anadkat, S. (2023). Gpt-4 Technical Report. arXiv."},{"key":"ref_32","unstructured":"Lambert, N., Morrison, J., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L.J.V., Liu, A., Dziri, N., and Lyu, S. (2024). Tulu 3: Pushing frontiers in open language model post-training. arXiv."},{"key":"ref_33","unstructured":"Lehmann, M. (2024). The Definitive Guide to Policy Gradients in Deep Reinforcement Learning: Theory, Algorithms and Implementations. arXiv, Available online: http:\/\/arxiv.org\/abs\/2401.13662."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Michailidis, P., Michailidis, I., and Kosmatopoulos, E. (2025). Reinforcement learning for optimizing renewable energy utilization in buildings: A review on applications and innovations. Energies, 18.","DOI":"10.3390\/en18071724"},{"key":"ref_35","unstructured":"Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., and Wu, Y. (2024). Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv."},{"key":"ref_36","unstructured":"Multi-Granularity, M.L.M.F. (2024, January 11\u201316). M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation. Proceedings of the ACL 2024, Bangkok, Thailand."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Muennighoff, N., Tazi, N., Magne, L., and Reimers, N. (2022). Mteb: Massive text embedding benchmark. arXiv.","DOI":"10.18653\/v1\/2023.eacl-main.148"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., and Shleifer, S. (2023). PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel. arXiv, Available online: http:\/\/arxiv.org\/abs\/2304.11277.","DOI":"10.14778\/3611540.3611569"},{"key":"ref_39","unstructured":"Kim, T., and Vossen, P. (2021). EmoBERTa: Speaker-Aware Emotion Recognition in Conversation with RoBERTa. arXiv, Available online: http:\/\/arxiv.org\/abs\/2108.12009."},{"key":"ref_40","unstructured":"Xiao, S., Liu, Z., Zhang, P., and Muennighoff, N. (2023). C-Pack: Packaged Resources To Advance General Chinese Embedding. arXiv, Available online: http:\/\/arxiv.org\/abs\/2309.07597."},{"key":"ref_41","unstructured":"Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., and Huang, F. (2024). Qwen2 Technical Report. arXiv, Available online: http:\/\/arxiv.org\/abs\/2407.10671."},{"key":"ref_42","unstructured":"Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., and Bhosale, S. (2023). Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv, Available online: http:\/\/arxiv.org\/abs\/2307.09288."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Bai, G., Liu, J., Bu, X., He, Y., Liu, J., Zhou, Z., Lin, Z., Su, W., Ge, T., and Zheng, B. (2024). Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues. arXiv.","DOI":"10.18653\/v1\/2024.acl-long.401"},{"key":"ref_44","unstructured":"Cai, Z., Cao, M., Chen, H., Chen, K., Chen, K., Chen, X., Chen, X., Chen, Z., Chen, Z., and Chu, P. (2024). InternLM2 Technical Report. arXiv, Available online: http:\/\/arxiv.org\/abs\/2403.17297."},{"key":"ref_45","unstructured":"Jiang, A.Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D.S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., and Saulnier, L. (2023). Mistral 7B. arXiv, Available online: http:\/\/arxiv.org\/abs\/2310.06825."},{"key":"ref_46","unstructured":"Li, A., Gong, B., Yang, B., Shan, B., Liu, C., Zhu, C., Zhang, C., Guo, C., Chen, D., and Li, D. (2025). Minimax-01: Scaling foundation models with lightning attention. arXiv."},{"key":"ref_47","unstructured":"GLM, T., Zeng, A., Xu, B., Wang, B., Zhang, C., Yin, D., Zhang, D., Rojas, D., Feng, G., and Zhao, H. (2024). Chatglm: A family of large language models from glm-130b to glm-4 all tools. arXiv."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/9\/738\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:33:33Z","timestamp":1760034813000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/16\/9\/738"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,27]]},"references-count":47,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2025,9]]}},"alternative-id":["info16090738"],"URL":"https:\/\/doi.org\/10.3390\/info16090738","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,27]]}}}