{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T00:30:16Z","timestamp":1760056216720,"version":"build-2065373602"},"reference-count":25,"publisher":"World Scientific Pub Co Pte Ltd","issue":"10","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272201"],"award-info":[{"award-number":["62272201"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Soft. Eng. Knowl. Eng."],"published-print":{"date-parts":[[2025,10]]},"abstract":"<jats:p> In recent years, large language models (LLMs) have achieved remarkable progress in dialogue generation. However, existing automatic evaluation methods still face challenges in diverse scenarios, particularly in terms of limited generalization and low alignment with human judgment. To address these issues, we propose a paraphrase-based evaluation framework that integrates Abstract Meaning Representation (AMR) with general-purpose language models (GPT). This approach generates diverse paraphrases across lexical, syntactic and stylistic dimensions to enhance the coverage of traditional evaluation metrics. Experimental results show that incorporating paraphrase augmentation significantly improves the correlation between automatic metrics and human evaluation on multiple datasets. Additionally, extensive experiments on six mainstream LLMs demonstrate the effectiveness and generalizability of the proposed method. This study offers new insights into improving the human alignment of automatic evaluation and lays a foundation for the application and optimization of LLMs in open-domain dialogue systems. <\/jats:p>","DOI":"10.1142\/s0218194025500433","type":"journal-article","created":{"date-parts":[[2025,8,15]],"date-time":"2025-08-15T08:21:37Z","timestamp":1755246097000},"page":"1383-1398","source":"Crossref","is-referenced-by-count":0,"title":["Paraphrase-Augmented Evaluation for Dialogue Models: A Unified Framework with AMR and GPT-Based Rewriting"],"prefix":"10.1142","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-8347-3674","authenticated-orcid":false,"given":"Xiao","family":"Liu","sequence":"first","affiliation":[{"name":"School of Artificial Intelligence and Computer Science, Jiangnan University, 1800 Lihu Avenue, Wuxi 214122, Jiangsu, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9481-9599","authenticated-orcid":false,"given":"Zhenping","family":"Xie","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence and Computer Science, Jiangnan University, 1800 Lihu Avenue, Wuxi 214122, Jiangsu, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-4501-7190","authenticated-orcid":false,"given":"Juncheng","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence and Computer Science, Jiangnan University, 1800 Lihu Avenue, Wuxi 214122, Jiangsu, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9187-0591","authenticated-orcid":false,"given":"Senlin","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of Internet of Things Engineering, Wuxi Institute of Technology, Wuxi, Jiangsu, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2025,9,17]]},"reference":[{"key":"S0218194025500433BIB003","first-page":"2","volume":"55","author":"Sai A. B.","year":"2022","journal-title":"ACM Comput. Surv."},{"key":"S0218194025500433BIB004","first-page":"311","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics","author":"Papineni K.","year":"2002"},{"key":"S0218194025500433BIB005","first-page":"74","volume-title":"Text Summarization Branches Out","author":"Lin C.-Y.","year":"2004"},{"first-page":"1","volume-title":"Int. Conf. Learn. Represent (ICLR 2020)","author":"Zhang T.","key":"S0218194025500433BIB006"},{"key":"S0218194025500433BIB007","first-page":"27263","volume":"34","author":"Yuan W.","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"S0218194025500433BIB009","first-page":"65","volume-title":"Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and\/or Summarization","author":"Banerjee S.","year":"2005"},{"key":"S0218194025500433BIB011","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11321"},{"key":"S0218194025500433BIB012","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.552"},{"key":"S0218194025500433BIB013","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.131"},{"key":"S0218194025500433BIB014","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.nlp4convai-1.5"},{"key":"S0218194025500433BIB015","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-long.365"},{"key":"S0218194025500433BIB018","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1009"},{"key":"S0218194025500433BIB019","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1042"},{"key":"S0218194025500433BIB020","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00318"},{"key":"S0218194025500433BIB022","doi-asserted-by":"publisher","DOI":"10.5121\/csit.2024.140418"},{"key":"S0218194025500433BIB025","series-title":"Part of ACL 2013 workshop series","first-page":"178","volume-title":"Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse","volume":"7","author":"Banarescu L.","year":"2013"},{"key":"S0218194025500433BIB026","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.57"},{"issue":"3","key":"S0218194025500433BIB028","first-page":"85","volume":"3","author":"AlAfnan M. A.","year":"2023","journal-title":"J. Artif. Intell. Technol."},{"key":"S0218194025500433BIB029","first-page":"76","volume-title":"Proceedings of the 1st Workshop on Personalization of Generative AI Systems (PERSONALIZE 2024)","author":"Bhandarkar A.","year":"2024"},{"key":"S0218194025500433BIB031","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i14.17489"},{"key":"S0218194025500433BIB032","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-acl.244"},{"key":"S0218194025500433BIB034","first-page":"312","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations and ACL 2020","author":"Goodman M. W.","year":"2020"},{"key":"S0218194025500433BIB035","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.nlp4convai-1.20"},{"key":"S0218194025500433BIB036","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2018.09.004"},{"key":"S0218194025500433BIB037","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.64"}],"container-title":["International Journal of Software Engineering and Knowledge Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0218194025500433","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T09:09:36Z","timestamp":1760000976000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/10.1142\/S0218194025500433"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,17]]},"references-count":25,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2025,10]]}},"alternative-id":["10.1142\/S0218194025500433"],"URL":"https:\/\/doi.org\/10.1142\/s0218194025500433","relation":{},"ISSN":["0218-1940","1793-6403"],"issn-type":[{"type":"print","value":"0218-1940"},{"type":"electronic","value":"1793-6403"}],"subject":[],"published":{"date-parts":[[2025,9,17]]}}}