{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T22:53:55Z","timestamp":1781823235627,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":25,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,12,22]],"date-time":"2023-12-22T00:00:00Z","timestamp":1703203200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,12,22]]},"DOI":"10.1145\/3639631.3639658","type":"proceedings-article","created":{"date-parts":[[2024,2,16]],"date-time":"2024-02-16T06:08:33Z","timestamp":1708063713000},"page":"158-163","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Large Language Models Evaluate Machine Translation via Polishing"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-2583-9235","authenticated-orcid":false,"given":"Yiheng","family":"Wang","sequence":"first","affiliation":[{"name":"Peking University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,2,16]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Proceedings of the Second Workshop on Statistical Machine Translation, Chris Callison-Burch, Philipp Koehn, Cameron\u00a0Shaw Fordyce","author":"Callison-Burch Chris","unstructured":"Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2007. (Meta-) Evaluation of Machine Translation. In Proceedings of the Second Workshop on Statistical Machine Translation, Chris Callison-Burch, Philipp Koehn, Cameron\u00a0Shaw Fordyce, and Christof Monz (Eds.). Association for Computational Linguistics, Prague, Czech Republic, 136\u2013158. https:\/\/aclanthology.org\/W07-0718"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324919000469"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.870"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.747"},{"key":"e_1_3_2_1_5_1","volume-title":"\u00a0T. Martins","author":"Fomicheva Marina","year":"2022","unstructured":"Marina Fomicheva, Shuo Sun, Erick Fonseca, Chrysoula Zerva, Fr\u00e9d\u00e9ric Blain, Vishrav Chaudhary, Francisco Guzm\u00e1n, Nina Lopatina, Lucia Specia, and Andr\u00e9 F.\u00a0T. Martins. 2022. MLQE-PE: A Multilingual Quality Estimation and Post-Editing Dataset. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, Nicoletta Calzolari, Fr\u00e9d\u00e9ric B\u00e9chet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, H\u00e9l\u00e8ne Mazo, Jan Odijk, and Stelios Piperidis (Eds.). European Language Resources Association, Marseille, France, 4963\u20134974. https:\/\/aclanthology.org\/2022.lrec-1.530"},{"key":"e_1_3_2_1_6_1","unstructured":"Mingqi Gao Jie Ruan Renliang Sun Xunjian Yin Shiping Yang and Xiaojun Wan. 2023. Human-like Summarization Evaluation with ChatGPT. arxiv:2304.02554\u00a0[cs.CL]"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1017\/S1351324915000339"},{"key":"e_1_3_2_1_8_1","unstructured":"Lifeng Han. 2022. An Overview on Machine Translation Evaluation. arxiv:2202.11027\u00a0[cs.CL]"},{"key":"e_1_3_2_1_9_1","unstructured":"Tom Kocmi and Christian Federmann. 2023. Large Language Models Are State-of-the-Art Evaluators of Translation Quality. In Proceedings of the 24th Annual Conference of the European Association for Machine Translation Mary Nurminen Judith Brenner Maarit Koponen Sirkku Latomaa Mikhail Mikhailov Frederike Schierl Tharindu Ranasinghe Eva Vanmassenhove Sergi\u00a0Alvarez Vidal Nora Aranberri Mara Nunziatini Carla\u00a0Parra Escart\u00edn Mikel Forcada Maja Popovic Carolina Scarton and Helena Moniz (Eds.). European Association for Machine Translation Tampere Finland 193\u2013203. https:\/\/aclanthology.org\/2023.eamt-1.19"},{"key":"e_1_3_2_1_10_1","volume-title":"Proceedings of the 11th Conference of the Association for Machine Translation in the Americas, Sharon O\u2019Brien","author":"Lacruz Isabel","year":"2014","unstructured":"Isabel Lacruz, Michael Denkowski, and Alon Lavie. 2014. Cognitive demand and cognitive effort in post-editing. In Proceedings of the 11th Conference of the Association for Machine Translation in the Americas, Sharon O\u2019Brien, Michel Simard, and Lucia Specia (Eds.). Association for Machine Translation in the Americas, Vancouver, Canada, 73\u201384. https:\/\/aclanthology.org\/2014.amta-wptp.6"},{"key":"e_1_3_2_1_11_1","volume-title":"ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74\u201381. https:\/\/aclanthology.org\/W04-1013"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Yang Liu Dan Iter Yichong Xu Shuohang Wang Ruochen Xu and Chenguang Zhu. 2023. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. arxiv:2303.16634\u00a0[cs.CL]","DOI":"10.18653\/v1\/2023.emnlp-main.153"},{"key":"e_1_3_2_1_13_1","unstructured":"Qingyu Lu Baopu Qiu Liang Ding Kanjian Zhang Tom Kocmi and Dacheng Tao. 2023. Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models: A Case Study on ChatGPT. arxiv:2303.13809\u00a0[cs.CL]"},{"key":"e_1_3_2_1_14_1","volume-title":"Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL","author":"Nenkova Ani","year":"2004","unstructured":"Ani Nenkova and Rebecca Passonneau. 2004. Evaluating Content Selection in Summarization: The Pyramid Method. In Proceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL 2004. Association for Computational Linguistics, Boston, Massachusetts, USA, 145\u2013152. https:\/\/aclanthology.org\/N04-1019"},{"key":"e_1_3_2_1_15_1","volume-title":"Advances in Neural Information Processing Systems, S.\u00a0Koyejo, S.\u00a0Mohamed, A.\u00a0Agarwal, D.\u00a0Belgrave, K.\u00a0Cho, and A.\u00a0Oh (Eds.). Vol.\u00a035. Curran Associates","author":"Ouyang Long","year":"2022","unstructured":"Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul\u00a0F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, S.\u00a0Koyejo, S.\u00a0Mohamed, A.\u00a0Agarwal, D.\u00a0Belgrave, K.\u00a0Cho, and A.\u00a0Oh (Eds.). Vol.\u00a035. Curran Associates, Inc., 27730\u201327744. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2022\/file\/b1efde53be364a73914f58805a001731-Paper-Conference.pdf"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-91241-7_7"},{"key":"e_1_3_2_1_18_1","volume-title":"Proceedings of the Sixth Workshop on Statistical Machine Translation, Chris Callison-Burch, Philipp Koehn, Christof Monz, and Omar\u00a0F. Zaidan (Eds.). Association for Computational Linguistics","author":"Popovi\u0107 Maja","year":"2011","unstructured":"Maja Popovi\u0107, David Vilar, Eleftherios Avramidis, and Aljoscha Burchardt. 2011. Evaluation without references: IBM1 scores as evaluation metrics. In Proceedings of the Sixth Workshop on Statistical Machine Translation, Chris Callison-Burch, Philipp Koehn, Christof Monz, and Omar\u00a0F. Zaidan (Eds.). Association for Computational Linguistics, Edinburgh, Scotland, 99\u2013103. https:\/\/aclanthology.org\/W11-2109"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.213"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.704"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Chenhui Shen Liying Cheng Xuan-Phi Nguyen Yang You and Lidong Bing. 2023. Large Language Models are Not Yet Human-Level Evaluators for Abstractive Summarization. arxiv:2305.13091\u00a0[cs.CL]","DOI":"10.18653\/v1\/2023.findings-emnlp.278"},{"key":"e_1_3_2_1_22_1","volume-title":"Proceedings of the 9th Conference of the Association for Machine Translation in the Americas: Research Papers. Association for Machine Translation in the Americas","author":"Specia Lucia","year":"2010","unstructured":"Lucia Specia and Jes\u00fas Gim\u00e9nez. 2010. Combining Confidence Estimation and Reference-based Metrics for Segment-level MT Evaluation. In Proceedings of the 9th Conference of the Association for Machine Translation in the Americas: Research Papers. Association for Machine Translation in the Americas, Denver, Colorado, USA. https:\/\/aclanthology.org\/2010.amta-papers.3"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"crossref","unstructured":"Jiaan Wang Yunlong Liang Fandong Meng Zengkui Sun Haoxiang Shi Zhixu Li Jinan Xu Jianfeng Qu and Jie Zhou. 2023. Is ChatGPT a Good NLG Evaluator? A Preliminary Study. arxiv:2303.04048\u00a0[cs.CL]","DOI":"10.18653\/v1\/2023.newsum-1.1"},{"key":"e_1_3_2_1_24_1","volume-title":"Natural Language Processing and Chinese Computing, Fei Liu, Nan Duan, Qingting Xu, and Yu\u00a0Hong (Eds.)","author":"Wu Ning","unstructured":"Ning Wu, Ming Gong, Linjun Shou, Shining Liang, and Daxin Jiang. 2023. Large Language Models are Diverse Role-Players for\u00a0Summarization Evaluation. In Natural Language Processing and Chinese Computing, Fei Liu, Nan Duan, Qingting Xu, and Yu\u00a0Hong (Eds.). Springer Nature Switzerland, Cham, 695\u2013707."},{"key":"e_1_3_2_1_25_1","volume-title":"Target-Side Language Model for\u00a0Reference-Free Machine Translation Evaluation","author":"Zhang Min","unstructured":"Min Zhang, Xiaosong Qiao, Hao Yang, Shimin Tao, Yanqing Zhao, Yinlu Li, Chang Su, Minghan Wang, Jiaxin Guo, Yilun Liu, and Ying Qin. 2022. Target-Side Language Model for\u00a0Reference-Free Machine Translation Evaluation. In Machine Translation, Tong Xiao and Juan Pino (Eds.). Springer Nature Singapore, Singapore, 45\u201353."}],"event":{"name":"ACAI 2023: 2023 6th International Conference on Algorithms, Computing and Artificial Intelligence","location":"Sanya China","acronym":"ACAI 2023"},"container-title":["2023 6th International Conference on Algorithms Computing and Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639631.3639658","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3639631.3639658","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,29]],"date-time":"2025-08-29T17:39:31Z","timestamp":1756489171000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639631.3639658"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,22]]},"references-count":25,"alternative-id":["10.1145\/3639631.3639658","10.1145\/3639631"],"URL":"https:\/\/doi.org\/10.1145\/3639631.3639658","relation":{},"subject":[],"published":{"date-parts":[[2023,12,22]]},"assertion":[{"value":"2024-02-16","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}